Luna 2 | See The Total Cost of Your Evals At Scale

Inputs

Agent archetype

Evaluator model

More models

Custom Edit rates

Number of agents

Traces per agent / day

Baseline shared metrics

Custom metrics per agent

Metric scope

Annual savings with Luna

LLM-as-judge costs 27× more than Luna at this scale

Per-eval response latency

Luna is faster per evaluation

Tokens scanned / mo

Trace volume

LLM-as-judge cost

Luna cost

Same budget. Different coverage.

If you pin the budget at Luna's cost ($146.1K/mo)

Blind spot with judge at equal budget

Annual eval cost vs. # agents

# AGENTS ANNUAL COST
LLM-as-judge Luna

Methodology & sources

Parameter Notes Source
Simple · 3K tok Single-turn or short RAG. Anchored to ~3,700 tok/ticket. Anthropic support agent
Tool-Using · 10K tok ReAct loop, real tool surface, 3–8 iterations. τ-bench · τ²-bench leaderboard
Code/Research · 50K tok Multi-step code or browsing. Claude Code 33K, Cursor 188K on SWE-bench. Cognition SWE-bench · HAL · GAIA
Judge rates Pick a model in the inputs panel. Effective rate uses 90% input / 10% output blend. Anthropic · OpenAI
F1 accuracy hit Smaller judges sacrifice eval accuracy. Mini-class −5% F1, Nano-class −10%. Patronus LLM-Judge Leaderboard
Luna 2 rate Luna 2 $0.12/1M (self-hosted). Fine-tuned per eval task. Galileo Observability
Per-eval latency Luna 2 ≈ 80ms p50; frontier judges 2.5–4s. Luna 2 paper
Metric scope Pooled: quadratic in agents. Per-agent: linear. Modeling choice