Code Judges
Score your evals with a custom Python function
A code judge scores your eval with a Python function you write, instead of a model or a built-in check. It's the escape hatch for rules that are real and checkable, but too complex for the other judge types: scoring against a lookup table, parsing a structured output and validating it, or measuring something specific to your domain.
Code judges are fast and cheap compared to LLM as Judge, and unlike an LLM judge they return the same score every time.
Creating a Code Judge
Add a judge to your eval and select the "Code" type.
Code judges run on your machine, with full access.
There are no import restrictions or resource limits beyond the timeout. Adding or editing code in a project requires confirming you trust it — that trust gate is the security boundary. See Code Trust for details.
The score() Function
Your code must define a function named score(). Kiln passes it only the arguments your function actually declares, so declare the ones you need and omit the rest. You must accept at least output or trace.
output
The model's final output, as a string
trace
The full conversation, as a list of message dicts
task_input
The original task input, as a string
reference_data
A dict of ground-truth data for this item, or None
score() returns a dict, keyed by the JSON key of each of your eval's output scores. A score's JSON key is its display name in lowercase snake_case: a score named "Exact Match" has the key exact_match. Kiln checks the returned keys against exactly this set, so return the JSON keys, not the display names.
def score(output: str) -> dict:
"""Pass when the output is valid JSON with a non-empty summary."""
import json
try:
parsed = json.loads(output)
except json.JSONDecodeError:
return {"valid_output": 0.0}
has_summary = bool(parsed.get("summary", "").strip())
return {"valid_output": 1.0 if has_summary else 0.0}Scores are floats. For a pass/fail score, return 1.0 or 0.0.
Calling LLMs from a Code Judge
Code judges can call two built-in tools, selectable from the judge's tool picker like any other tool:
llm: a general-purpose model call. Pass a prompt, model and provider. Provide an optional JSON schema to force structured output; without one you get text back.llm_judge: the same, but it automatically applies your eval's own output score schema, so the scores come back keyed the way your eval expects.
Both take a prompt that is rendered as a Jinja2 template against an input dict. Put your instructions in the template and pass the run's data through input — never build the prompt by interpolating trace text directly, or a {{ }} appearing in that text will be parsed as Jinja and fail the call.
Both run in Kiln itself rather than in your judge's subprocess, so your API keys are never exposed to your code.
This combination solves the long-trace problem. An agent run can produce a 500k token trace, and handing all of it to an LLM judge is slow and expensive. Instead, filter it down in code to the handful of messages that actually matter, then ask a cheap model about just those:
Calling Other Tools
The LLM tools aren't special: a code judge can call any tool in its allowlist, using the same from kiln import tools API that code tools use. That lets a judge look up ground truth in your own systems, or reuse a tool you've already written.
See Calling Other Tools for the full API, including async calls and error handling.
Timeout and Tools
Timeout: wall-clock timeout for scoring one item, including any nested tool calls. Defaults to 180 seconds, with a 300 second maximum.
Tools: an explicit allowlist of the tools your code may call. Code with no tools selected can't make tool calls at all.
Your Code is a Real Python File
Kiln stores your judge's source as scorer.py, in the eval config's folder next to eval_config.kiln:
Because it's a plain Python file rather than a string inside JSON, it's importable, lintable, type-checkable, and produces a readable diff when you review changes in Git.
Testing with pytest
You can write standard pytest tests against your judge. score() takes plain keyword arguments and depends only on the standard library plus kiln_ai, so you can import and call it directly — no Kiln-specific test runner.
Create test_scorer.py beside scorer.py:
Then run pytest from that folder. Kiln doesn't store, display, or run these tests — they live in a normal Python environment with kiln_ai installed.
Learn More
Judge Types: all the judge types, and when to use each
Code Tools: the same Python authoring model, for tools your agents call
Code Tools Authoring Guide: the complete authoring contract, including the tool calling API and testing details
Last updated