DATASETBENCHMARKS
AGENTTESTING

ResearchlabfocusedonrubricsforAItraining.

We study and design evaluation rubrics to supplement frontier AI labs and agent builders.

Rubric library

Versioned, peer-reviewed criteria definitions.

Benchmark packs

Curated test scenarios ready for agent evaluation.

Sandbox environments

Controlled environments for safe workflow testing.

Read Blog
Scroll

On this page

01 / 08

Rubric Library

Human domain expert creators, editors and reviewers work alongside assisted automation to expand our rubric repository built from realistic workflows — structured representations of judgment that define what "good" looks like for AI systems and agents.

Hard scenariosPrompts and rubricsVersionedReusablePeer-reviewedExecutable

Rule-based

Deterministic pass/fail criteria

LLM-judged

Model-assisted quality assessment

Hybrid

Combined human + automated evaluation

Used for:Benchmark creationSandbox evaluationResearch studiesAgent performance measurementAI training
rubric.expert
Prompt
WritePreview

Maddox Reyes has been championing the full-time employee onboarding process for Vireon Health Insurance. While Senna's onboarding went well, some teammates hit significant access issues.

Find the most recent Vireon welcome email and note the support contact details. Find the most recent Health Insurance event Maddox organized, capture the meeting link, organizer email, and Senna's email, then compare it with the invitation to confirm the year.

Find who told Maddox they had portal access problems, their names, issues, and emails. Reply to that thread with the quickest resolution path, then schedule \u201cHealth Insurance Team onboarding support\u201d for July 4, 2026 at 1 PM for 30 minutes, with Senna as organizer and the captured meeting link as location.

App Environments
ABC
Rubric Itemsonboarding · gmail retrieval
5
#1

Gmail was opened.

gmailnavigation
5
#2

“Vireon” was the term that was entered into the search bar.

gmailemail_searchretrieval
5
#3

The search result list was accessed.

gmaildetail_searchretrieval
10
#4

The most recent Vireon welcome email was identified with the subject “Welcome to Vireon! Check your benefits Now”.

gmaildetail_searchretrieval
10
#5

The date of the most recent Vireon welcome email was identified as January 20th, 2025.

gmaildetail_searchretrievaltemporal_reasoning
02 / 08

Archetypes

Extraction

Identify relevant information from inputs

Synthesis

Combine information across sources correctly

Decision-Making

Choose correct next step in workflow

Push Action

Perform correct tool or system actions

Communication

Produce correct, safe, appropriate outputs

03 / 08

Benchmark Packs

Curated test scenarios ready for agent evaluation. Each pack includes prompts, expected outcomes, and scoring rubrics — run across multiple models to compare performance.

83% pass rate
rubric.expert
3 passed
1 borderline
1 failed
40 scenarios
BP-01

Billing dispute resolution

GPT-4ClaudeGemini
92%
2026-06-14
BP-02

Multi-step contract review

GPT-4Claude
87%
2026-06-14
BP-03

Support ticket triage

GPT-4ClaudeGeminiMistral
76%
2026-06-13
BP-04

Calendar conflict resolution

GPT-4Claude
94%
2026-06-15
BP-05

Invoice reconciliation

GPT-4Claude
68%
2026-06-12
04 / 08

Sandbox Environments

UI applications are engineered into software sandbox environments — controlled, reproducible settings where agents are tested against real-world workflows.

Each sandbox includes

Initial state definition
Available tools & actions
Observations & feedback loops
Ground truth outcomes
Rubric-based evaluation layer

Example Scenario

Handling a billing dispute on an enterprise account

A long-standing enterprise customer disputes a charge, claiming their contract includes a grandfathered rate. The agent must investigate across billing, contracts, and communication history, then resolve within compliance guidelines. The agent must:

1Read and classify the dispute from the support ticket
2Retrieve the customer's contract and billing history
3Cross-reference the disputed charge against contract terms
4Review past communications for any rate agreements
5Calculate the correct amount or confirm the charge is valid
6Draft a professional resolution response with findings

Measures: Extraction accuracy · Cross-reference precision · Policy compliance · Calculation correctness · Communication tone · Resolution completeness

Gmailsandbox
Slacksandbox
Zendesksandbox
UI applications engineered into software sandbox environments
05 / 08

Why a Research Lab

Standards are emerging

No universal benchmark exists. Evaluation standards are being defined now.

Empirical refinement

Rubrics are continuously tested against real agent behaviour.

Infrastructure output

Research becomes reusable infrastructure for evaluation pipelines.

06 / 08

How to Plug In

Two ways to use rubric.expert — request our curated rubric datasets for AI training, or connect your agent to our MCP server for live evaluation inside simulated app environments.

Use case

AI Training

Request our curated rubric datasets — prompts, scenarios, expected outcomes, and grading rubrics — to train or fine-tune your models.

Curated prompt → rubric pairs
Human-verified expected outcomes
Structured evaluation criteria
Cross-domain scenario coverage
Request dataset →

Use case

Evaluation

Connect your agent to our MCP server and test it inside simulated app environments — scored against our rubrics.

Connect via MCP server
Your agent calls simulated tool APIs
Each action scored by rubric
Receive full evaluation report
MCP endpoint →
rubric.expert MCP
Your AgentOpenAI / Anthropic / etc.
MCP connect
MCP Servermcp.rubric.expert
routes to
A
App Sim 1
B
App Sim 2
C
App Sim 3

Available tool calls via MCP

search_records()read_item()get_entry()send_message()post_update()modify_entry()query_index()create_draft()
07 / 08

Continuous Improvement

New rubrics, new benchmarks, and more app environments added in updates.

Phase I

Foundation

Core rubric library covering the most common agent workflows across fintech, support, and legal domains.

Phase II

Pattern Recognition

Cross-domain evaluation patterns identified, abstracted, and standardised into reusable templates.

Phase III

Standardisation

Industry-wide benchmark standardisation. Rubrics become the common language for agent quality.

Phase IV

The Reference Layer

The de facto reference for agent evaluation. Every model ships with a rubric.expert score.

08 / 08

Pricing

Compare your build versus buy.

Build Internally

Domain expert workshops£10,000–£50,000
Benchmark design£5,000–£25,000
Sandbox engineering£20,000–£100,000+
Ongoing maintenanceContinuous
TotalSignificant ongoing investment

Rubric.Expert

Workflow benchmark packFixed license
Sandbox accessSubscription
UpdatesIncluded
Time to deploymentImmediate

Bring your own agent. Run realistic workflows. Measure performance against expert-reviewed standards.

Contact us for benchmark pack and sandbox pricing — hi@rubric.expert

Logo on white

Mark / On White

Logo on black

Mark / On Black

rubric.expert

Wordmark / Light

rubric.expert

Wordmark / Dark

About Us

rubric.expert is built by Windowshop AI, a company dedicated to building infrastructure for AI evaluation and agent testing.

Learn about our team

Test agents in controlled evaluation environments.

Explore benchmark packs, sandbox environments, and rubric research.

Read Blog

hi@rubric.expert

rubric.expert

Research lab focused on rubrics for AI training.

Explore

Blog

Company

Brandhi@rubric.expert

Legal

© 2026 rubric.expert. All rights reserved.Rubrics. Benchmarks. Sandbox environments. Controlled evaluation systems.

rubric.expert

By using rubric.expert, you accept our