How Abstractus performs at full-text screening on 3,338 public papers
Today, we are talking about our evaluation of our full text screening. Much like our previous abstract screening post, similar metrics and evaluations will be used.
Data availability
By far the most important issue to address first is data availability. Currently, anyone seeking to perform easy mass retrieval of PDFs requires the use of open access websites and API links. Therefore, while the original abstract dataset contained 50,000+ abstracts across 18 studies, only about 3,338 PDFs could be retrieved for screening and our dataset size was reduced to 11.
Another issue that arises is the relatively small number of included (true positive) papers in these small subsets. We of course excluded those datasets without any papers labeled as included post-retrieval, but thought it was still reasonable to analyze those that contained very few. This is reflective of the needle-in-the-haystack nature of most systematic review tasks, and is of course related to the fact that many reviews have a vanishingly small number of included papers to begin with.
How we measure performance
We opted to use recall once again, but decided to use specificity instead of workload reduction as one of our primary performance metrics. While on large datasets workload reduction works well to communicate information to more non-technical audiences, it is greatly affected by the overall number of papers. If, for example, there are 10 included PDFs in a set of 100 total, noting that our workload reduction figure was 90% would be uninformative, since the upper bound of this reduction is capped by the number of included papers.
This is why specificity is superior, because it measures how many excluded papers were correctly excluded. It is defined technically as TN / (TN + FP).
Recall is a measure of how many included papers were correctly included. It is defined technically as TP / (TP + FN).
Results
With all of that said, I will present our full text screening numbers, which came out fairly well. Our median specificity was 97% and our median recall was 100%. Our full benchmark results are below:
Recall and specificity by study
Full text screening on 11 SYNERGY studies, sorted by size. Median recall 100%, median specificity 97.3%.
- Recall
- Specificity
| Studies | 11 |
|---|---|
| Records screened | 3,338 |
| True positives | 127 |
| False positives | 216 |
| False negatives | 9 |
| True negatives | 2,986 |
| Median study weighted recall | 100.0% |
| Median study weighted precision | 44.4% |
| Median study weighted specificity | 97.3% |
| Median study weighted F1 | 0.615 |
| Study | TP | FP | FN | TN | N | Recall | Precision | Specificity | F1 |
|---|---|---|---|---|---|---|---|---|---|
| Donners_2021 | 3 | 3 | 0 | 15 | 21 | 100.0% | 50.0% | 83.3% | 0.667 |
| Hall_2012 | 2 | 2 | 0 | 72 | 76 | 100.0% | 50.0% | 97.3% | 0.667 |
| Leenaars_2019 | 2 | 2 | 0 | 314 | 318 | 100.0% | 50.0% | 99.4% | 0.667 |
| Leenaars_2020 | 77 | 138 | 6 | 429 | 650 | 92.8% | 35.8% | 75.7% | 0.517 |
| Menon_2022 | 9 | 17 | 0 | 107 | 133 | 100.0% | 34.6% | 86.3% | 0.514 |
| Sep_2021 | 14 | 0 | 0 | 48 | 62 | 100.0% | 100.0% | 100.0% | 1.000 |
| Wolters_2018 | 1 | 3 | 1 | 402 | 407 | 50.0% | 25.0% | 99.3% | 0.333 |
| van_Dis_2020 | 2 | 8 | 0 | 921 | 931 | 100.0% | 20.0% | 99.1% | 0.333 |
| van_de_Schoot_2018 | 4 | 5 | 0 | 309 | 318 | 100.0% | 44.4% | 98.4% | 0.615 |
| van_der_Valk_2021 | 9 | 5 | 2 | 81 | 97 | 81.8% | 64.3% | 94.2% | 0.720 |
| van_der_Waal_2022 | 4 | 33 | 0 | 288 | 325 | 100.0% | 10.8% | 89.7% | 0.195 |