CAS Bench
CAS Bench tests structured biotech asset sourcing under strict precision. On the 80 query held out test set, Convexia reached 0.921 R@P≥0.95, beating the best commercial baseline by 10.7 points, with 0.44 false positives per query. The benchmark measures whether a system can turn a thesis style sourcing query into qualifying drug assets with dated evidence. The primary metric is recall at precision ≥0.95 because false positives waste analyst time. Convexia leads on high precision recall, candidate precision and recall, hard queries, preclinical sourcing, China linked assets, and long tail source recovery.
Benchmark at a glance
The headline result and the design of the benchmark: a strict precision recall task over 80 held out queries and 4,137 gold assets, with a frozen November 1, 2025 evidence cutoff.
0.921
main benchmark score
+10.7
points vs Database A
0.969
returned asset precision
0.973
returned asset recall
0.44
analyst noise
4,137
gold asset universe
Primary leaderboard
Systems plotted on the source backed high precision recall frontier. Convexia sits on the upper right frontier with the lowest false positive rate per query.
High precision recall frontier
Primary leaderboard
Where the gap is largest
Convexia degrades least as query difficulty increases, and separates most on preclinical sourcing, the strongest product proof point.
Difficulty degradation
Development stage contrast
Long tail source moat
Recall by source type. The separation from commercial baselines is concentrated in long tail channels like TTO portfolios, theses, grants, and conference materials.
Source moat heatmap
China linked asset recovery
China linked gold assets are 1,218 of 4,137, equal to 29.4% of the CAS Bench test set. Convexia recovered 1,065, which is 309 more than the best commercial baseline.
China linked asset recovery
Reliability and analyst noise
Per query recall distribution and false positive burden. Convexia is both stronger and more consistent, keeping 55 of 80 queries at zero false positives.
Per query recall distribution
False positive triage grid
Ablation ladder
Each rung adds a capability layer. R@P rises and false positives fall. The advantage is not generic LLM browsing, it is retrieval coverage plus constraint aware normalization plus citation scoring.