Summary
| Stage | What we measure | Result | Reference set |
|---|---|---|---|
| Search | Among studies included in published reviews, how many the strategy proposed by revisia alone retrieves from PubMed | 91.8% (90.3–92.8% over three runs) | 45 reviews, 1,093 included studies |
| Title and abstract screening | How many included studies remain after screening with criteria suggested by the app | 99.4% (324 of 326; 95% CI 97.8–99.8). 24% of records go to human review | 16 reviews, 326 included studies, 3,526 records |
| Data extraction | Correct values, “not reported” when absent, and verbatim quotations from the text | 117 of 117 in the internal test, not yet checked by an independent person. In production, 98.7% of cells had a verified quotation | 39 questions across 8 articles; 1,359 cells from 30 days of real use |
All three measurements were made on 29 and 30 September 2026 using the code in production since 30 September. Your manuscript cites the platform version used for each stage, so figures here can be linked to the version used by your project.
How we measure in general
- Published reviews define the ground truth. We take the studies their authors included from each review's characteristics table and resolve them to PubMed identifiers. If the review included a study and revisia loses it, that counts as a failure.
- No manual corrections. The app receives the question and suggests the search strategy, criteria and decisions. We accept everything as a new user might. Our reviewers make no corrections along the way.
- Three runs. Language models do not answer identically twice. We repeat every measurement three times and publish the mean and range.
- The real product. The harness calls the same functions as the production app, with the same instructions and configuration for every plan.
- No model names. We do not publish which models or internal instructions each stage uses. They change, while the platform version date identifies a measurement. If your journal asks which exact model your project used, email equipo@revisia.es and we will provide it in writing.
- Continuous validation. Whenever search, criteria or screening changes, we measure again against the bank before deployment. If a figure no longer holds, we correct this page first. The review bank, including PubMed identifiers for included studies, scenarios and results for every run is stored in our repository with the date and measured version. Email equipo@revisia.es if you want a copy of the bank so you can repeat the check.
1. Search
What we measure. Recall of the strategy proposed by revisia: among studies included by a published review, the share returned by the PubMed query. This is the crucial search measure. Nobody can screen a study that the search never retrieves.
How. A bank of 45 open-access systematic reviews published from 2019 to 2026 across four question types: intervention (9), prevalence and incidence (12), diagnostic tests (12), and scoping reviews (12). A review enters only if at least five included studies can be found in PubMed, for 1,093 studies total. We enter each review question into the app, accept the suggested strategy, run it in PubMed without date limits and count how many of the 1,093 appear.
| Question type | Reviews | Included | Previous version (29 Sep, morning) | Current |
|---|---|---|---|---|
| Prevalence and incidence | 12 | 537 | 90.5% | 95.8% |
| Diagnostic tests | 12 | 227 | 73.4% | 93.7% |
| Intervention | 9 | 85 | 79.6% | 88.2% |
| Scoping reviews | 12 | 244 | 65.2% | 88.6% * |
| All | 45 | 1,093 | 80.5% | 91.8% (90.3–92.8) |
* Scoping reviews received a second improvement measured only on those 12 reviews, over three runs: 88.5%, 88.1% and 89.3%. The global 91.8% comes from the round before that improvement, when scoping reviews scored 82.2%. We have not rerun all 45 and do not combine the two figures.
- Half the reviews score 97.6% or higher. Nine out of ten studies are retrieved without anyone editing the query.
- Europe PMC, using the same translated strategy: 89.8%. It rescues no studies missed by PubMed because both queries fail for the same reason, a missing term. It currently acts as a fallback when PubMed does not respond, not a second chance.
- The cost: more records to screen. Median records per search rose from 1,881 to 5,032. One in four searches exceeds 34,000 and 17 of 45 exceed 10,000. In a project with a record allowance, some of that recall cannot be screened.
- Where it still fails: six reviews remain below 80%. In one review of exercise and anxiety in young people, the app does not suggest “yoga”, the intervention name used by several included studies, and recall is 62%. In another, a general word enters the wrong concept and the query returns 306,000 records. The quality checker flags it but no longer disables it, so the person must read the warning.
- Not measured: OpenAlex and ClinicalTrials.gov, which search also runs. The bank uses PubMed identifiers and cannot score them. Included items without a PubMed ID, such as theses, grey literature and unindexed journals, are outside the denominator. The figure means “of what PubMed contains”, not 100% of each review. The bank has only one Spanish question, which retrieved 6 of 6 studies in all three runs, so it says nothing about language.
2. Title and abstract screening
What we measure. Safety sensitivity: among studies included by a published review, the share that remain as “include” or “review” after screening. Doubt goes to a person, never into silence, so “review” counts as retained. We also measure the human workload and the share of non-included records that screening can discard.
How. Sixteen reviews from the same bank, four per question type, with 6 to 61 included studies each; one is in Spanish. None was used to tune search. Each corpus contains every included study with a PubMed ID plus 200 randomly selected non-included records returned by the app's query, limited to the review's publication year. The app proposes criteria from the question and search and we accept all of them. We do not use the published review's criteria. Three runs per corpus. Every plan uses the same screening configuration.
| Question type | Included studies retained | Sensitivity | Records sent to human review |
|---|---|---|---|
| Intervention | 58 of 58 | 100% | 17.6% |
| Prevalence and incidence | 117 of 118 | 99.2% | 33.0% |
| Diagnostic tests | 60 of 61 | 98.4% | 19.8% |
| Scoping reviews | 89 of 89 | 100% | 26.0% |
| All | 324 of 326 | 99.4% (95% CI 97.8–99.8) | 24.3% |
- The worst of three runs retains 323 of 326 (99.1%). No included study was lost because it lacked an abstract, was a protocol or secondary analysis, or because of participant age. None was lost through missing abstract data. The three cases below contain a verbatim abstract sentence that contradicts a criterion.
- The two stable losses are defensible from the abstract alone. One prevention drug trial does not name the diagnostic test studied by the review, so the design criterion excludes it with a quotation. The authors included it after reading full text. The other says participants were “recruited at outpatient clinics of nine hospitals” in a review requiring community studies. Its authors interpreted their own question more loosely.
- The third is an unstable reading error. In an enterocolitis-incidence review, one run read “enterocolitis (OR: 0.59)” as an associated factor rather than the frequency requested by the criterion. The other two runs sent it to review. We count it because a rule cannot fix this kind of failure, but a second reading can.
- Remaining human work: 24.3% of records go to “review”, ranging from 5.6% to 51.6% by review. Among sampled non-included records, screening excludes 67%, sends 25% to review and includes 8%. The 67% is a lower bound because the sample comes from the query and contains relevant studies the published review may never have seen.
- Non-determinism: 10.3% of records change decision across runs, usually between “include” and “review”. We therefore recommend checking a sample of proposed exclusions.
- In Spanish: three medical reviews with Spanish questions and criteria, on migraine, hypertension and ICU delirium, cover 337 records, 37 of them included. They retain 36 of 37 included studies across five runs, with 12.7% sent to review. The lost single-arm dose-titration study is excluded by comparator and design criteria with textual reasons; the original reviewers included it.
- Limits: the denominator contains included studies with PubMed identifiers, 326 of roughly 368 declared by the reviews. Non-included records are a random sample, not verified exclusions, so we do not publish specificity. Criteria come from one proposal per review, not a measurement of their variation. Only one of the 16 reviews is in Spanish.
3. Data extraction
What we measure. Four properties for every extracted item: whether the value is correct; whether it says “not reported” when the article does not provide it; whether the quoted sentence exists verbatim and contains the value; and whether repeated runs return the same answer.
How. Thirty-nine questions across eight real open-access full-text articles, 25 on Spanish articles and 14 on English articles. Every question was asked three times using the production pattern, one call per article with all its columns, for 117 answers.
| Measure | Result |
|---|---|
| Correct values | 117 of 117 |
| Wrong value presented as verified | 0 |
| Invented value when absent from the article | 0 |
| Correct “Not reported” | 15 of 15 |
| Quotations that exist verbatim in the text | 102 of 102 |
| Questions with the same answer in all three runs | 39 of 39 |
In production, during the 30 days before measurement, 1,341 of 1,359 cells with a value had a verified quotation (98.7%). The other 18 were marked “verify” for a person to inspect; none appeared without a quotation. “Verified” means the sentence exists verbatim in the article and contains the value, not that a person has checked whether the value is correct. These data cover only three internal-account projects, one providing 93% of all cells.
Why we do not claim “100% accuracy”. An AI agent wrote the 39 correct test answers during development while reading only the article, and no independent person has confirmed them. Six of 39 were revised after seeing model responses. With 117 answers, a perfect score does not rule out a true error rate of up to 3% per answer. Every text came from clean XML, without noisy PDFs. The next step is blind labelling by two clinically trained readers. We will publish it here with inter-rater agreement.
Why we do not publish a stopping rule
Prioritisation screening tools learn from your decisions and reorder the list so relevant items rise. They then need a rule that says when to stop reading. That rule creates both the saving and the risk: unread records are unseen, so a premature stop leaves no trace. Two 2026 studies quantified this. One evaluated fifteen stopping methods on eighty-one datasets and found that nearly all stop late or miss studies. The other repeated a published review of 4,690 records with a widely used prioritisation tool and retrieved 16 and 21 of the 30 included studies.
revisia avoids that problem because it does not use that design. We screen every record within the allowance with a criterion-level decision and send doubt to a person. The allowance itself is still a limit: records above it are marked “unscreened · human verification”, counted separately in PRISMA and left for you.
What we measured before
From June to August 2026, this page reported a smaller validation: four reviews and 94 included studies for screening, with pooled sensitivity of 96.8% and 27 of 27 retained in one Spanish review, plus a single-database search retrieving 7% to 45% depending on the corpus. Those figures describe product versions no longer running. The figures on this page replace them and use banks ten times larger. Previous results remain in our repository history with their dates.
Last updated: 2026-09-30