We're excited to announce that Omic has achieved the top score on the GPQA-Diamond benchmark with 93.3% accuracy—the highest ever recorded and a significant leap beyond both frontier AI models and human PhD experts.
What is GPQA-Diamond?
The Graduate-Level Google-Proof Q&A (GPQA) Diamond benchmark is considered one of the most challenging tests of scientific reasoning for AI systems. It consists of graduate-level questions in biology, chemistry, and physics that are designed to be:
- Expert-validated: Questions written and verified by domain PhD experts
- Google-proof: Answers cannot be easily found through web searches
- Genuinely difficult: Even human PhD experts in the relevant field achieve only ~65% accuracy
This makes GPQA-Diamond a true test of deep scientific understanding, not just information retrieval.
The Results
Our system achieved 93.3% accuracy, compared to:
| Model | Score |
|---|---|
| Omic | 93.3% |
| Gemini 3 Pro | 91.9% |
| GPT-5.1 | 88.1% |
| Grok 4 | 87.7% |
| Claude 4.5 Sonnet | 83.4% |
| Human PhD Expert | 65.0% |
This represents a 1.4 percentage point improvement over the next best system (Gemini 3 Pro) and a remarkable 28.3 percentage points above human PhD expert performance.
Why This Matters for Drug Discovery
GPQA-Diamond tests exactly the kind of reasoning required for pharmaceutical research:
- Molecular biology: Understanding gene regulation, protein function, and cellular mechanisms
- Chemistry: Predicting reaction outcomes, understanding molecular properties, and designing synthetic routes
- Physics: Modeling molecular dynamics, understanding thermodynamics, and predicting binding interactions
When an AI system can reason at this level across all three domains simultaneously, it becomes a powerful tool for drug discovery—connecting insights across disciplines in ways that accelerate the path from target to therapeutic.
Our Approach
Unlike general-purpose AI models, Omic's system is purpose-built for scientific reasoning in drug discovery:
- Specialized domain models: Fine-tuned on millions of scientific papers, patents, and experimental datasets
- Agentic reasoning: Multi-step problem decomposition with verification and self-correction
- Scientific tools: Integration with computational chemistry, molecular simulation, and biological databases
- Ensemble methods: Combining multiple reasoning pathways to increase accuracy and reduce errors
This isn't about building a better chatbot. It's about building AI that can do real science.
What's Next
This benchmark result validates our core thesis: AI systems optimized for scientific domains can dramatically outperform general-purpose models on the problems that matter for drug discovery.
We're now applying this capability across our pipeline programs, from target identification through lead optimization. The same reasoning that achieves 93.3% on graduate-level science questions is helping us discover novel therapeutics for oncology, infectious disease, cardiovascular disease, and hematology.
Stay tuned for more updates as we continue to push the boundaries of what AI can achieve in drug discovery.
Benchmark methodology: Accuracy measured on the full GPQA-Diamond test set using standard evaluation protocols. Human PhD baseline (65%) from the original GPQA paper. Competitor scores sourced from Artificial Analysis as of November 2025.
Omic - Building Biological Superintelligence to End Disease