We study and design evaluation rubrics to supplement frontier AI labs and agent builders.
Versioned, peer-reviewed criteria definitions.
Curated test scenarios ready for agent evaluation.
Controlled environments for safe workflow testing.
On this page
Human domain expert creators, editors and reviewers work alongside assisted automation to expand our rubric repository built from realistic workflows — structured representations of judgment that define what "good" looks like for AI systems and agents.
Deterministic pass/fail criteria
Model-assisted quality assessment
Combined human + automated evaluation
Maddox Reyes has been championing the full-time employee onboarding process for Vireon Health Insurance. While Senna's onboarding went well, some teammates hit significant access issues.
Find the most recent Vireon welcome email and note the support contact details. Find the most recent Health Insurance event Maddox organized, capture the meeting link, organizer email, and Senna's email, then compare it with the invitation to confirm the year.
Find who told Maddox they had portal access problems, their names, issues, and emails. Reply to that thread with the quickest resolution path, then schedule \u201cHealth Insurance Team onboarding support\u201d for July 4, 2026 at 1 PM for 30 minutes, with Senna as organizer and the captured meeting link as location.
Gmail was opened.
“Vireon” was the term that was entered into the search bar.
The search result list was accessed.
The most recent Vireon welcome email was identified with the subject “Welcome to Vireon! Check your benefits Now”.
The date of the most recent Vireon welcome email was identified as January 20th, 2025.
Identify relevant information from inputs
Combine information across sources correctly
Choose correct next step in workflow
Perform correct tool or system actions
Produce correct, safe, appropriate outputs
Curated test scenarios ready for agent evaluation. Each pack includes prompts, expected outcomes, and scoring rubrics — run across multiple models to compare performance.
Billing dispute resolution
Multi-step contract review
Support ticket triage
Calendar conflict resolution
Invoice reconciliation
UI applications are engineered into software sandbox environments — controlled, reproducible settings where agents are tested against real-world workflows.
Each sandbox includes
Example Scenario
A long-standing enterprise customer disputes a charge, claiming their contract includes a grandfathered rate. The agent must investigate across billing, contracts, and communication history, then resolve within compliance guidelines. The agent must:
Measures: Extraction accuracy · Cross-reference precision · Policy compliance · Calculation correctness · Communication tone · Resolution completeness
No universal benchmark exists. Evaluation standards are being defined now.
Rubrics are continuously tested against real agent behaviour.
Research becomes reusable infrastructure for evaluation pipelines.
Two ways to use rubric.expert — request our curated rubric datasets for AI training, or connect your agent to our MCP server for live evaluation inside simulated app environments.
Use case
Request our curated rubric datasets — prompts, scenarios, expected outcomes, and grading rubrics — to train or fine-tune your models.
Use case
Connect your agent to our MCP server and test it inside simulated app environments — scored against our rubrics.
Available tool calls via MCP
New rubrics, new benchmarks, and more app environments added in updates.
Phase I
Core rubric library covering the most common agent workflows across fintech, support, and legal domains.
Phase II
Cross-domain evaluation patterns identified, abstracted, and standardised into reusable templates.
Phase III
Industry-wide benchmark standardisation. Rubrics become the common language for agent quality.
Phase IV
The de facto reference for agent evaluation. Every model ships with a rubric.expert score.
Compare your build versus buy.
Build Internally
Rubric.Expert
Bring your own agent. Run realistic workflows. Measure performance against expert-reviewed standards.
Contact us for benchmark pack and sandbox pricing — hi@rubric.expert

Mark / On White

Mark / On Black
rubric.expertWordmark / Light
rubric.expertWordmark / Dark
rubric.expert is built by Windowshop AI, a company dedicated to building infrastructure for AI evaluation and agent testing.
Learn about our teamExplore benchmark packs, sandbox environments, and rubric research.