Luna Evaluation Models Cloud Observability | Splunk

Splunk Agent Observability

Luna Evaluation Models

Purpose-built small language models that evaluate and guardrail every interaction, so you can watch 100% of traffic affordably.

Free edition - Try it free for 14 days — no credit card required.

Take a guided tour - Got 5 minutes? See how it works.

98% cheaper changes the economics

When a verdict costs about $0.15 per million tokens instead of $5.00, you stop rationing. Grade and guard every interaction, not a sample.

Coverage

Cost

Accuracy

The LLM-as-judge tax

Luna is built for one job, judging agent interactions. That single focus is why it can grade and guard 100% of traffic in real time, where a frontier judge is too expensive and too slow.

Too expensive to grade everything

A frontier judge can cost around $5.00 per million scanned tokens. At production volume that is too expensive to run on everything, so teams check 1 to 5% of interactions and hope the rest look the same. The ones you miss are the ones that hurt: the hallucination you never saw, the injection that slipped through.

Too slow to stop anything

A frontier judge can take about 3,200ms to return a verdict. By then the tool has already run, so the verdict can only describe what happened. It can audit, but it cannot guard. Luna returns a verdict in about 152ms, fast enough to block an action before the tool fires.

Tune Luna to your domain, no code required

Out-of-the-box metrics get you started, but your domain has its own definition of right. With Luna Studio, correct a handful of Luna's verdicts and it learns your standard, reaching around 95% accuracy on your own tasks. No labeling pipeline, no model training expertise. Review, correct, and ship a tuned metric, all from the studio.

Take judges from training to production

Tune a judge in Luna Studio, and Splunk runs it in production for you. No servers to set up, no MLOps team required. Tune a judge for each check that matters, ready to score and guard in real time.

Scale real-time evaluations without scaling cost

Explore the documentation

Evaluate every interaction in production

Run always-on evaluations without the cost and latency of larger models. Luna-2’s fine-tuned SLMs deliver millisecond-level verdicts at pennies per million tokens, making high-scale production evaluation practical.

Catch risky agent actions before they execute

Guardrail agentic workflows before mistakes become actions. Luna-2 evaluates tool selection and agent flows in real time, catching risky behavior before tools execute rather than only detecting problems afterward.

Run more checks without slowing agents down

Evaluate multiple dimensions of agent behavior at once without adding seconds of delay. Luna-2 can run 10–20 checks simultaneously in under 200 milliseconds on L4 GPUs, enabling real-time guardrails at scale.

Cut the cost of production evaluations

Get production-grade evaluation without paying for a frontier LLM judge on every interaction. Luna-2 delivers evaluations at $0.12 per million tokens while maintaining high accuracy and low latency.

Know whether agents actually complete the job

Measure more than the final response. Luna-2 evaluates tool errors, tool selection quality, action advancement, and action completion so you can see whether agents successfully move toward and accomplish user goals.

Scale the metrics that matter to your application

Evaluate the behaviors specific to your AI application. Luna-2 can power custom LLM metrics for production use cases, while lightweight adapters allow one base model to scale across hundreds of metrics with minimal infrastructure overhead.

Luna Evaluation Models FAQs

What is Luna?

Luna is a family of purpose-built small language models for evaluation. They grade and guard agent interactions quickly and cheaply, which is what makes it affordable to watch 100% of traffic instead of a sample.

Why not use a frontier LLM as the judge?

A frontier judge is expensive and slow, so teams sample 1 to 5% of traffic and can only review actions after they run. Luna is built for one job, judging agent interactions, so it is cheap enough to grade everything and fast enough to guard in real time.

How much cheaper is Luna than an LLM judge?

Luna runs evaluations and guardrails at roughly 98% lower cost than an LLM-as-judge approach, which changes the economics from rationing to running on every interaction.

Is Luna scoring repeatable?

Yes. Luna returns a verdict as a single token, so scores are deterministic: run the same input twice and you get the same answer, which is what makes a model trustworthy as a judge.

Can Luna be tuned without code?

Yes. With Luna Studio you correct a handful of verdicts and it learns your standard, reaching around 95% accuracy on your own tasks, with no labeling pipeline or model-training expertise.