← All posts

How Abstractus performs at full-text screening on 3,338 public papers

  • benchmarks
  • methodology
A grid of full papers laid out for screening, with the few that meet the criteria shown in orange

Today, we are talking about our evaluation of our full text screening. Much like our previous abstract screening post, similar metrics and evaluations will be used.

Data availability

By far the most important issue to address first is data availability. Currently, anyone seeking to perform easy mass retrieval of PDFs requires the use of open access websites and API links. Therefore, while the original abstract dataset contained 50,000+ abstracts across 18 studies, only about 3,338 PDFs could be retrieved for screening and our dataset size was reduced to 11.

Another issue that arises is the relatively small number of included (true positive) papers in these small subsets. We of course excluded those datasets without any papers labeled as included post-retrieval, but thought it was still reasonable to analyze those that contained very few. This is reflective of the needle-in-the-haystack nature of most systematic review tasks, and is of course related to the fact that many reviews have a vanishingly small number of included papers to begin with.

How we measure performance

We opted to use recall once again, but decided to use specificity instead of workload reduction as one of our primary performance metrics. While on large datasets workload reduction works well to communicate information to more non-technical audiences, it is greatly affected by the overall number of papers. If, for example, there are 10 included PDFs in a set of 100 total, noting that our workload reduction figure was 90% would be uninformative, since the upper bound of this reduction is capped by the number of included papers.

This is why specificity is superior, because it measures how many excluded papers were correctly excluded. It is defined technically as TN / (TN + FP).

Recall is a measure of how many included papers were correctly included. It is defined technically as TP / (TP + FN).

Results

With all of that said, I will present our full text screening numbers, which came out fairly well. Our median specificity was 97% and our median recall was 100%. Our full benchmark results are below:

Recall and specificity by study

Full text screening on 11 SYNERGY studies, sorted by size. Median recall 100%, median specificity 97.3%.

  • Recall
  • Specificity
Overall totals across all 11 studies.
Studies11
Records screened3,338
True positives127
False positives216
False negatives9
True negatives2,986
Median study weighted recall100.0%
Median study weighted precision44.4%
Median study weighted specificity97.3%
Median study weighted F10.615
Per-study results. Study names match the SYNERGY dataset identifiers.
StudyTPFPFNTNNRecallPrecisionSpecificityF1
Donners_20213301521100.0%50.0%83.3%0.667
Hall_20122207276100.0%50.0%97.3%0.667
Leenaars_2019220314318100.0%50.0%99.4%0.667
Leenaars_202077138642965092.8%35.8%75.7%0.517
Menon_20229170107133100.0%34.6%86.3%0.514
Sep_202114004862100.0%100.0%100.0%1.000
Wolters_201813140240750.0%25.0%99.3%0.333
van_Dis_2020280921931100.0%20.0%99.1%0.333
van_de_Schoot_2018450309318100.0%44.4%98.4%0.615
van_der_Valk_2021952819781.8%64.3%94.2%0.720
van_der_Waal_20224330288325100.0%10.8%89.7%0.195