The standard for AI on systems of recordThe standard for measuring AI on enterprise systems-of-record
SOR-Bench Workday TextQA
SOR-Bench evaluates which AI models can correctly answer questions and take actions in enterprise systems-of-record like Workday, ServiceNow, Salesforce, and NetSuite.
- Integrations
- Security
- Reporting
- People Experience (PEX)
- Core HCM
- Studio
- Extend
- Payroll
- Prism Analytics
- Recruiting
- Revenue Management
- Student
- Talent & Performance
- Absence
- Learning
- Financial Planning
- Financial Management (FIN)
- Accounting
- Compensation
- Adaptive Planning
- Benefits
- Procurement
- Advanced Comp
- Time Tracking
- Expenses
- Peakon Employee Voice
- VNDLY
- Platform
- Financial Reporting
- Scheduling
- Data Conversion
- Integrations
- Security
- Reporting
- People Experience (PEX)
- Core HCM
- Studio
- Extend
- Payroll
- Prism Analytics
- Recruiting
- Revenue Management
- Student
- Talent & Performance
- Absence
- Learning
- Financial Planning
- Financial Management (FIN)
- Accounting
- Compensation
- Adaptive Planning
- Benefits
- Procurement
- Advanced Comp
- Time Tracking
- Expenses
- Peakon Employee Voice
- VNDLY
- Platform
- Financial Reporting
- Scheduling
- Data Conversion
Hover over a module to preview a sample question
What's missing
Existing AI benchmarks don't measure performance on mission-critical enterprise applications
They tell you how AI models fare on taking hypothetical tests. SOR-Bench tells you how they fare on real system integration challenges.
SOR-Bench evaluates AI performance over the following enterprise applications
- Live

- Coming soon

- Coming soon

- Coming soon

- Coming soon

- Coming soon

- Coming soon

- Coming soon

Methodology
How SOR-Bench works
What we test, how we score, and how to submit your system.
Corpus
Real administrative scenarios across configuration, troubleshooting, policy interpretation, and release impact.
Coverage
Every major module and administrative domain in the system of record. No cherry-picked subset.
Difficulty
Easy, Medium, and Hard tiers calibrated to administrative workload, not surface complexity.
Scoring
Every answer is judged against an expert-verified solution on correctness, citation quality, safety, and actionability.
Evidence
Citations must be real, current, and traceable to a primary source. Ungrounded fluency does not score.
Safety
Hallucinations and risky administrative recommendations are flagged and penalized independently.
Held-out
The scoring corpus and expert-verified solutions stay private. Public samples are synthetic and representative.
Submissions
Vendors and labs submit models or agents through a controlled evaluation pipeline.
Compete
Think your AI can beat SOR-Bench?
Submit a model or agent for evaluation. Public leaderboard or confidential. Your call.
