How Abstractus performs on a public benchmark of 57,177 abstracts
Our abstract screening model, run on a public, university-designed benchmark of over 57,000 abstracts, achieves a median recall of 95% and reduces the number of abstracts a researcher would need to screen by a median of 84%.
Hello again!
We wanted to share today why you should trust Abstractus. In an age where automated processes are increasingly used, it is vitally important that systems performing critical research operations are able to work reliably with clear validation.
We are therefore happy to share both the original data and the outcomes for our abstract screening pipeline. This allows you to check for yourself how our model performs, rather than merely having to rely upon our private benchmarking. This is made possible by researchers at Utrecht University in the Netherlands, who released the SYNERGY dataset in 2023. It is a repository of abstracts, titles, and the final decisions research teams made pertaining to their inclusion in their systematic reviews.
Why a public benchmark matters
Our use of this dataset is relevant because, usually, studies benchmarking the performance of automation software on systematic reviews depend on private, custom-curated datasets, often acquired by asking select authors for data from their systematic review. They are typically only created for one study and then are thrown away or are inaccessible. This is why we advocate for the use and creation of new, quality public benchmarks in this research space.
We want to be clear that the path many researchers have used for their dataset selection is not necessarily their fault, for data accessibility is notoriously difficult given the proprietary nature of so many research articles, but we do hope this is a situation that improves in the future.
Which studies we used
We used 19 of the 26 studies found in the SYNERGY dataset, including studies based upon methodological and quality assessments. As is standard in data processing, we also removed any rows that were missing information, leaving us with 57,177 abstracts. The seven removed studies were excluded for the following reasons:
- No dual screening for the entire review — Moran 2021, Wassenaar 2017, Smid 2020.
- Title screening only — Muthu 2021.
- Quality issues — Jeyaraman 2020.
- Very large studies, exceeding 25,000 papers — Brouwer 2019, Walker 2018.
The first two problem areas are fairly explicable. An abnormal methodological procedure would not be a fair comparison for any automated system, including ours. We seek to use studies with a “gold standard” design, which always involves two screeners making decisions using both the title and abstract.
However, problem areas 3 and 4 warrant further explanation, as these decision areas are less obvious. Quality issues are more subjective than screening methodologies, but, in general, issues like mismatches in language between the studies’ described objectives and the data searched would be one glaring example. To be specific, in Jeyaraman 2020, the study is supposed to be on stem cells and spinal cord injury, yet the keywords used during the literature search phase were “Knee Osteoarthritis,” “Knee Degeneration,” “Stem Cell Therapy,” “Mesenchymal Stem Cells,” and “Bone Marrow.” While some of these search terms related to the original question (namely those talking about stem cells), the lack of any keywords pertaining to the spine indicates a clear and obvious mismatch. This, along with other quality issues, was relevant to this study’s exclusion.
Problem area 4, involving study size, was a practical decision made out of concerns over spending efficiency and meaningful research insights gained. Screening nearly 80,000 extra abstracts for two studies involving about 40,000 abstracts each would add very little overall understanding of how effective an automated pipeline is, not to mention causing substantially larger research fees by more than doubling the number of documents included in the study. Even more importantly, our ability to run multiple tests and experiments on this dataset would have been greatly impeded by longer waiting times between results.
How we measure performance
When conducting our research, we used two primary metrics: recall and workload reduction.
Recall (also known as sensitivity) is defined technically as TP / (TP + FN), where TP stands for true positive and FN for false negative. In our specific case this measures how many studies that ended up in the final review were also included by our system. A recall of 100% would mean we found all of the relevant studies. A recall of 80% would mean we missed 20% of them. This is by far the most important metric for any automated literature review software, as missing essential studies would critically impact the main conclusions of a reviewer using any tool prospectively.
We also opted to use recall over a metric like accuracy, given the highly imbalanced nature of literature reviews. Out of hundreds of potential records, only a small number of studies end in the final analysis. This means that any tool that simply guesses “exclude” for all of the studies in a dataset would have accuracy near 99%.
The second metric is workload reduction. We define this as the number of papers labeled as “exclude” by a system. A workload reduction of 50% would mean that half of the studies were found to be irrelevant and excluded. This metric is important for an automated system since it measures how meaningfully researchers would be assisted.
Workload reduction is constrained by a few factors. The first is the necessity of erring on the side of caution when including studies, as our modeling approach is tuned to expansively include studies given the importance of recall. The second is the absolute size of the candidate set of abstracts. Generally, smaller abstract sets tend to have much more narrowly defined keyword search terms and therefore a higher relevance per paper, leading to a higher inclusion rate and limiting the potential workload reduction.
Results
| Studies | 19 |
|---|---|
| Records screened | 57,177 |
| True positives | 1,124 |
| False positives | 6,567 |
| False negatives | 71 |
| True negatives | 49,415 |
| Median study weighted recall | 95.0% |
| Median study weighted workload reduction | 84.0% |
| Study | TP | FP | FN | TN | N | Recall | Precision | Specificity | F1 | Workload reduction |
|---|---|---|---|---|---|---|---|---|---|---|
| Appenzeller-Herzog_2019 | 25 | 135 | 0 | 1,653 | 1,813 | 100.0% | 15.6% | 92.4% | 0.270 | 91.2% |
| Bos_2018 | 9 | 132 | 1 | 4,301 | 4,443 | 90.0% | 6.4% | 97.0% | 0.119 | 96.8% |
| Chou_2003 | 12 | 46 | 1 | 1,438 | 1,497 | 92.3% | 20.7% | 96.9% | 0.338 | 96.1% |
| Chou_2004 | 7 | 199 | 2 | 1,060 | 1,268 | 77.8% | 3.4% | 84.2% | 0.065 | 83.8% |
| Donners_2021 | 13 | 97 | 1 | 136 | 247 | 92.9% | 11.8% | 58.4% | 0.210 | 55.5% |
| Hall_2012 | 99 | 108 | 5 | 8,468 | 8,680 | 95.2% | 47.8% | 98.7% | 0.637 | 97.6% |
| Leenaars_2019 | 15 | 840 | 1 | 4,600 | 5,456 | 93.8% | 1.8% | 84.6% | 0.034 | 84.3% |
| Leenaars_2020 | 450 | 1,212 | 12 | 3,983 | 5,657 | 97.4% | 27.1% | 76.7% | 0.424 | 70.6% |
| Meijboom_2021 | 34 | 138 | 1 | 620 | 793 | 97.1% | 19.8% | 81.8% | 0.329 | 78.3% |
| Menon_2022 | 71 | 134 | 3 | 759 | 967 | 95.9% | 34.6% | 85.0% | 0.509 | 78.8% |
| Nelson_2002 | 70 | 104 | 7 | 143 | 324 | 90.9% | 40.2% | 57.9% | 0.558 | 46.3% |
| Oud_2018 | 19 | 66 | 1 | 814 | 900 | 95.0% | 22.4% | 92.5% | 0.362 | 90.6% |
| Radjenovic_2013 | 48 | 466 | 0 | 5,357 | 5,871 | 100.0% | 9.3% | 92.0% | 0.171 | 91.2% |
| Sep_2021 | 39 | 43 | 1 | 187 | 270 | 97.5% | 47.6% | 81.3% | 0.639 | 69.6% |
| Wolters_2018 | 14 | 25 | 4 | 3,594 | 3,637 | 77.8% | 35.9% | 99.3% | 0.491 | 98.9% |
| van_Dis_2020 | 70 | 1,755 | 1 | 6,733 | 8,559 | 98.6% | 3.8% | 79.3% | 0.074 | 78.7% |
| van_de_Schoot_2018 | 35 | 371 | 3 | 3,817 | 4,226 | 92.1% | 8.6% | 91.1% | 0.158 | 90.4% |
| van_der_Valk_2021 | 62 | 67 | 26 | 514 | 669 | 70.5% | 48.1% | 88.5% | 0.571 | 80.7% |
| van_der_Waal_2022 | 32 | 629 | 1 | 1,238 | 1,900 | 97.0% | 4.8% | 66.3% | 0.092 | 65.2% |
What this means in practice
As shown above, our median recall is 95.0% and our median workload reduction is 84%. This means researchers can consistently automate their initial screening entirely, or at the very least massively reduce the remaining papers they need to screen manually, all while finding the essential papers their study needs.
However, these quality metrics become more meaningful when the rapidity of this process is noted. Abstractus can screen a typical set of 2,000 abstracts in 10 to 15 minutes. The value of this becomes immediately apparent to any researcher who has started a screening and then realized they needed to refine their criteria. Typically this would require them to talk to other people on their screening team, receive approval for their new criteria, and then restart the process from the beginning. This means losing weeks or months worth of work.
No doubt a non-negligible number of reviewers become aware of substantial flaws in their screening process, but because of cost, time, and pressure to publish, they never restart.
Such practices impact not only the quality of their own research but also those who rely upon meta-analyses to make informed decisions, especially in high-risk fields like medicine. Now, with Abstractus, a researcher could run the exact same abstract screening 4 to 5 times within an hour and choose the one they like the most. This both makes correcting mistakes easier and allows teams to iteratively refine their searches as a core part of their methodological approach, choosing the screening results they believe best answer their core research question.
We hope you found this article informative and that you have a better understanding of Abstractus, our methodology, our validation, and our value to researchers. If you are interested in accessing our proven capabilities, feel free to create an account and sign up for a subscription that best suits your needs. If you have further questions or would like to get in touch, you can use our contact page.