Principles, not a frozen rubric
How benchmark results are documented and interpreted
A useful benchmark needs more than a score: it needs a clear task scope, a reviewable basis for assessment, and enough provenance to understand what a result does and does not say.
Current status: No approved benchmark results have been published. This page describes planned scope and publication principles, not a completed study. The general principles here are not a finalized universal rubric. Each published run's recorded rubric version is fixed provenance for that run.
01 · Task definition
Publish the task scope with each dataset
Enterprise answers depend on context. A task may rely on a particular business object, configuration, security boundary, workflow state, or release. A future dataset should state the assumptions needed to interpret its questions and identify its version, rather than imply that one set of tasks represents every tenant.
Initial task areas:
These descriptions define planned topic scope, not a claim that every category's question set is complete. Any published results are listed separately by run below.
02 · Provenance
Keep runs reproducible and distinct
Each published cohort identifies its dataset and rubric versions, product version (or explicitly says when unknown), publication date, evaluation track, question count, and run environment and tools. A model-only run and an agent-assisted run are different evaluation settings; the declared track and environment make that distinction legible.
Different dataset, rubric, product, or execution versions can change what is being measured. Runs remain separate public snapshots; ERPBenchmarks does not combine unlike cohorts into a site-wide score or rank.
03 · Assessment
Review before publication
The planned review standard is for tasks and the assessment approach to be reviewed by people with relevant domain knowledge before a run is published. The project owner reports experience across Workday® Finance, Reporting, Integrations, and Adaptive Planning; that experience is not presented as a certification, product endorsement, or substitute for transparent review of a particular run.
The rubric version listed for an approved run is fixed provenance for that run. It does not make the general principles on this page a finalized or shared scoring rubric; criteria for future runs remain subject to review and versioning.
04 · Coverage and failures
Make denominators visible
Question count, completed responses, assessed responses, and attempts describe different parts of a run. Published results should retain those counts so readers can see the coverage behind a measure. A failed, unanswered, or unassessed task should not disappear merely because it is difficult to score; any reason for exclusion needs to be explicit in the published methodology.
A completion rate is not the same as assessed quality. A score should be interpreted only alongside its coverage and the versioned rubric that produced it.
05 · Separate measures
Do not collapse quality and efficiency
Quality, completion, latency, and cost answer different questions and should remain separate. No score formula, weighting, or efficiency threshold is finalized here. An unavailable measurement should be reported as unknown rather than written as zero; a measured zero remains a real value and must not be confused with missing data.
Timing and cost can vary with model version, provider settings, tools, and the execution environment. A result should describe those conditions and avoid treating an efficiency measurement as evidence of answer quality.
What is public today
Planned coverage, no published outcomes
The public site currently provides planned category scope and publication principles only. No approved benchmark results are published. Category descriptions do not claim that questions are complete or that product behavior has been verified.
Return to the category overview