Benchmark reportFrozen evidence cutoff November 1, 2025

CAS Bench

CAS Bench tests structured biotech asset sourcing under strict precision. On the 80 query held out test set, Convexia reached 0.921 R@P≥0.95, beating the best commercial baseline by 10.7 points, with 0.44 false positives per query. The benchmark measures whether a system can turn a thesis style sourcing query into qualifying drug assets with dated evidence. The primary metric is recall at precision ≥0.95 because false positives waste analyst time. Convexia leads on high precision recall, candidate precision and recall, hard queries, preclinical sourcing, China linked assets, and long tail source recovery.

01

Benchmark at a glance

The headline result and the design of the benchmark: a strict precision recall task over 80 held out queries and 4,137 gold assets, with a frozen November 1, 2025 evidence cutoff.

PrimaryHero metrics · Convexia on the 80 query test set
R@P≥0.95

0.921

main benchmark score

Gain vs best baseline

+10.7

points vs Database A

Candidate precision

0.969

returned asset precision

Candidate recall

0.973

returned asset recall

FP / query

0.44

analyst noise

Test gold assets

4,137

gold asset universe

02

Primary leaderboard

Systems plotted on the source backed high precision recall frontier. Convexia sits on the upper right frontier with the lowest false positive rate per query.

Figure 1

High precision recall frontier

Figure 2

Primary leaderboard

03

Where the gap is largest

Convexia degrades least as query difficulty increases, and separates most on preclinical sourcing, the strongest product proof point.

Figure 3

Difficulty degradation

Figure 4

Development stage contrast

04

Long tail source moat

Recall by source type. The separation from commercial baselines is concentrated in long tail channels like TTO portfolios, theses, grants, and conference materials.

Figure 5

Source moat heatmap

05

China linked asset recovery

China linked gold assets are 1,218 of 4,137, equal to 29.4% of the CAS Bench test set. Convexia recovered 1,065, which is 309 more than the best commercial baseline.

Figure 6

China linked asset recovery

06

Reliability and analyst noise

Per query recall distribution and false positive burden. Convexia is both stronger and more consistent, keeping 55 of 80 queries at zero false positives.

Figure 7

Per query recall distribution

Figure 8

False positive triage grid

07

Ablation ladder

Each rung adds a capability layer. R@P rises and false positives fall. The advantage is not generic LLM browsing, it is retrieval coverage plus constraint aware normalization plus citation scoring.

Figure 9

Ablation ladder