Luna Evaluation Models Cloud Observability | Splunk
Splunk Agent Observability
Luna Evaluation Models
Purpose-built small language models that evaluate and guardrail every interaction, so you can watch 100% of traffic affordably.
Free edition - Try it free for 14 days — no credit card required.
Take a guided tour - Got 5 minutes? See how it works.
98% cheaper changes the economics
When a verdict costs about $0.15 per million tokens instead of $5.00, you stop rationing. Grade and guard every interaction, not a sample.
Coverage
- 172% of traffic graded, not a sample
Cost
- 169% cheaper than LLM-as-judge
Accuracy
- 163 on par with a frontier judge
The LLM-as-judge tax
Luna is built for one job, judging agent interactions. That single focus is why it can grade and guard 100% of traffic in real time, where a frontier judge is too expensive and too slow.
Too expensive to grade everything
A frontier judge can cost around $5.00 per million scanned tokens. At production volume that is too expensive to run on everything, so teams check 1 to 5% of interactions and hope the rest look the same. The ones you miss are the ones that hurt: the hallucination you never saw, the injection that slipped through.
Too slow to stop anything
A frontier judge can take about 3,200ms to return a verdict. By then the tool has already run, so the verdict can only describe what happened. It can audit, but it cannot guard. Luna returns a verdict in about 152ms, fast enough to block an action before the tool fires.
Tune Luna to your domain, no code required
Out-of-the-box metrics get you started, but your domain has its own definition of right. With Luna Studio, correct a handful of Luna's verdicts and it learns your standard, reaching around 95% accuracy on your own tasks. No labeling pipeline, no model training expertise. Review, correct, and ship a tuned metric, all from the studio.
Take judges from training to production
Tune a judge in Luna Studio, and Splunk runs it in production for you. No servers to set up, no MLOps team required. Tune a judge for each check that matters, ready to score and guard in real time.
Scale real-time evaluations without scaling cost
Evaluate every interaction in production
Run always-on evaluations without the cost and latency of larger models. Luna-2’s fine-tuned SLMs deliver millisecond-level verdicts at pennies per million tokens, making high-scale production evaluation practical.
Catch risky agent actions before they execute
Guardrail agentic workflows before mistakes become actions. Luna-2 evaluates tool selection and agent flows in real time, catching risky behavior before tools execute rather than only detecting problems afterward.
Run more checks without slowing agents down
Evaluate multiple dimensions of agent behavior at once without adding seconds of delay. Luna-2 can run 10–20 checks simultaneously in under 200 milliseconds on L4 GPUs, enabling real-time guardrails at scale.
Cut the cost of production evaluations
Get production-grade evaluation without paying for a frontier LLM judge on every interaction. Luna-2 delivers evaluations at $0.12 per million tokens while maintaining high accuracy and low latency.
Know whether agents actually complete the job
Measure more than the final response. Luna-2 evaluates tool errors, tool selection quality, action advancement, and action completion so you can see whether agents successfully move toward and accomplish user goals.
Scale the metrics that matter to your application
Evaluate the behaviors specific to your AI application. Luna-2 can power custom LLM metrics for production use cases, while lightweight adapters allow one base model to scale across hundreds of metrics with minimal infrastructure overhead.
Luna Evaluation Models FAQs
What is Luna?
Luna is a family of purpose-built small language models for evaluation. They grade and guard agent interactions quickly and cheaply, which is what makes it affordable to watch 100% of traffic instead of a sample.
Why not use a frontier LLM as the judge?
A frontier judge is expensive and slow, so teams sample 1 to 5% of traffic and can only review actions after they run. Luna is built for one job, judging agent interactions, so it is cheap enough to grade everything and fast enough to guard in real time.
How much cheaper is Luna than an LLM judge?
Luna runs evaluations and guardrails at roughly 98% lower cost than an LLM-as-judge approach, which changes the economics from rationing to running on every interaction.
Is Luna scoring repeatable?
Yes. Luna returns a verdict as a single token, so scores are deterministic: run the same input twice and you get the same answer, which is what makes a model trustworthy as a judge.
Can Luna be tuned without code?
Yes. With Luna Studio you correct a handful of verdicts and it learns your standard, reaching around 95% accuracy on your own tasks, with no labeling pipeline or model-training expertise.