01 · Task definition

Publish the task scope with each dataset

Enterprise answers depend on context. A task may rely on a particular business object, configuration, security boundary, workflow state, or release. A future dataset should state the assumptions needed to interpret its questions and identify its version, rather than imply that one set of tasks represents every tenant.

Initial task areas:

These descriptions define planned topic scope, not a claim that every category's question set is complete. Any published results are listed separately by run below.

02 · Provenance

Keep runs reproducible and distinct

Each published cohort identifies its dataset and rubric versions, product version (or explicitly says when unknown), publication date, evaluation track, question count, and run environment and tools. A model-only run and an agent-assisted run are different evaluation settings; the declared track and environment make that distinction legible.

Different dataset, rubric, product, or execution versions can change what is being measured. Runs remain separate public snapshots; ERPBenchmarks does not combine unlike cohorts into a site-wide score or rank.

03 · Assessment

Review before publication

The planned review standard is for tasks and the assessment approach to be reviewed by people with relevant domain knowledge before a run is published. The project owner reports experience across Workday® Finance, Reporting, Integrations, and Adaptive Planning; that experience is not presented as a certification, product endorsement, or substitute for transparent review of a particular run.

The rubric version listed for an approved run is fixed provenance for that run. It does not make the general principles on this page a finalized or shared scoring rubric; criteria for future runs remain subject to review and versioning.

04 · Coverage and failures

Make denominators visible

Question count, completed responses, assessed responses, and attempts describe different parts of a run. Published results should retain those counts so readers can see the coverage behind a measure. A failed, unanswered, or unassessed task should not disappear merely because it is difficult to score; any reason for exclusion needs to be explicit in the published methodology.

A completion rate is not the same as assessed quality. A score should be interpreted only alongside its coverage and the versioned rubric that produced it.

05 · Separate measures

Do not collapse quality and efficiency

Quality, completion, latency, and cost answer different questions and should remain separate. No score formula, weighting, or efficiency threshold is finalized here. An unavailable measurement should be reported as unknown rather than written as zero; a measured zero remains a real value and must not be confused with missing data.

Timing and cost can vary with model version, provider settings, tools, and the execution environment. A result should describe those conditions and avoid treating an efficiency measurement as evidence of answer quality.

What is public today

Planned coverage, no published outcomes

The public site currently provides planned category scope and publication principles only. No approved benchmark results are published. Category descriptions do not claim that questions are complete or that product behavior has been verified.

Return to the category overview