Whitepaper
Get the SOR-Bench Workday TextQA whitepaper
How today's AI actually performs on real Workday administration.
Key findings
What you'll get
A lab-grade breakdown of how four AI systems handle real Workday administration — scored by module and difficulty.
Scored results across every module
How each system performed by Workday module and difficulty, fully visualized.
Where AI fails on Workday
The failure patterns we saw on troubleshooting, configuration, and policy questions.
How it was scored
The grading approach and calibration behind every number on the leaderboard.
What it means for your team
How to read these results as you automate documentation, troubleshooting, and testing.
Why the questions stay private
A score only means something if the system never saw the answers ahead of time. So SOR-Bench works like a professional certification exam: the questions and the official answer key stay sealed, and every system is tested on questions it has never seen.
That is what makes a result reflect real Workday ability instead of memorization. The whitepaper shares the observations and the data, not the private questions or their answers.
Read the full methodology