# Welcome to Kiln AI

Rapid AI Prototyping and Dataset Collaboration Tool

Kiln AI is the easiest tool for:

* 🎛️ [**Fine Tuning**](/docs/fine-tuning/fine-tuning-guide): Zero-code fine-tuning for Llama, GPT-4o, and more. Automatic serverless deployment of models.
* 📊 [**Evals**](/docs/evals-and-specs/evaluations): Evaluate the quality of your models/tasks using state of the art evaluators.
* 📈 [**Optimizers**](/docs/optimizers): Find the best model, prompt or fine-tune for your task. Improve quality and lower costs.
* 🔍 [**Docs & Search (RAG)**](/docs/documents-and-search-rag): Add knowledge to your AI systems with Retrieval-Augmented Generation (RAG)
* 🤖 [**Agents**](/docs/agents): Build agentic systems with multiple actors
* 🛠 [**Tools & MCP**](/docs/tools-and-mcp): Connect powerful tools to your Kiln tasks, or write your own in Python with [Code Tools](/docs/tools-and-mcp/code-tools)
* 🧩 [**Skills**](/docs/skills): Load more instructions, depending on the goal.
* 🪄 [**Synthetic Data Generation**](/docs/synthetic-data-generation): Generate training data with our interactive visual tooling.
* 🧠 [**Reasoning Models**](/docs/fine-tuning/guide-train-a-reasoning-model): Train or distill your own custom reasoning models.
* 🤝 [**Team Collaboration**](/docs/collaboration): Git-based version control for your AI datasets. Intuitive UI makes it easy to collaborate with QA, PM, and subject matter experts on structured data (examples, prompts, ratings, feedback, issues, etc.).

### Jump right in

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden></th><th data-hidden data-card-target data-type="content-ref"></th><th data-hidden data-card-cover data-type="image">Cover image</th></tr></thead><tbody><tr><td><strong>Quickstart</strong></td><td>Download our app, setup a task, and run AI models.</td><td></td><td><a href="/pages/CyH2xJQs9yWJ1S8BYNav">/pages/CyH2xJQs9yWJ1S8BYNav</a></td><td><a href="/files/dOIQCelxqLNy37Lbe1FW">/files/dOIQCelxqLNy37Lbe1FW</a></td></tr><tr><td><strong>Automatic Prompt Optimizer</strong></td><td>Find the best possible prompt for your task.</td><td></td><td><a href="/pages/3Zl4LTmOhIK5etcBO7Yw">/pages/3Zl4LTmOhIK5etcBO7Yw</a></td><td><a href="/files/CYsusrOr631L4LlBM60z">/files/CYsusrOr631L4LlBM60z</a></td></tr><tr><td><strong>Models and AI Providers</strong></td><td>Manage which models you can use in Kiln.</td><td></td><td><a href="/pages/DCVyepbQxO6TlZjCkseN">/pages/DCVyepbQxO6TlZjCkseN</a></td><td><a href="/files/PZdTyzmLfT96PTT2iODf">/files/PZdTyzmLfT96PTT2iODf</a></td></tr><tr><td><strong>Fine Tuning</strong></td><td>Our guide for fine tuning AI models</td><td></td><td><a href="/pages/UQ5mUIArq32hihhGE4Sv">/pages/UQ5mUIArq32hihhGE4Sv</a></td><td><a href="/files/reaAySjXgTjdVkcX2Tzy">/files/reaAySjXgTjdVkcX2Tzy</a></td></tr><tr><td><strong>Evals</strong></td><td>Evaluate the quality of your models/tasks</td><td></td><td><a href="/pages/EJPf7zdXy05scsLk0wbf">/pages/EJPf7zdXy05scsLk0wbf</a></td><td><a href="/files/CvkRKvaj4K8rDekvQqEd">/files/CvkRKvaj4K8rDekvQqEd</a></td></tr><tr><td><strong>Tools &#x26; MCP</strong></td><td>Connect tools to Kiln</td><td></td><td><a href="/pages/qmzsaLmyoO8DcirfJpAK">/pages/qmzsaLmyoO8DcirfJpAK</a></td><td><a href="/files/RqmwElvFlPnm4q3vJ1Nq">/files/RqmwElvFlPnm4q3vJ1Nq</a></td></tr><tr><td><strong>Code Tools</strong></td><td>Write a Python function that runs as a tool</td><td></td><td><a href="/pages/3NrIuXJC45iAC1tRDfkf">/pages/3NrIuXJC45iAC1tRDfkf</a></td><td><a href="/files/DrxYj7KvQVl0GRHKPQea">/files/DrxYj7KvQVl0GRHKPQea</a></td></tr><tr><td><strong>Search Tools / RAG</strong></td><td>Add documents and create search indexes</td><td></td><td><a href="/pages/Dtsf7EZfJgPB9VbSUiae">/pages/Dtsf7EZfJgPB9VbSUiae</a></td><td><a href="/files/opZURxUeL4URf4LA99i5">/files/opZURxUeL4URf4LA99i5</a></td></tr><tr><td><strong>Agents</strong></td><td>Build multi-action agentic systems</td><td></td><td><a href="/pages/6X2jCKYVfnoJLkVMp41g">/pages/6X2jCKYVfnoJLkVMp41g</a></td><td><a href="/files/HUfoCrZmOkhj4AOFybMJ">/files/HUfoCrZmOkhj4AOFybMJ</a></td></tr><tr><td><strong>Synthetic Data Generation</strong></td><td>Generate training data for models with AI</td><td></td><td><a href="/pages/uwLdgH3kTWx6g0u1ca9R">/pages/uwLdgH3kTWx6g0u1ca9R</a></td><td><a href="/files/DpTjun7gVZipSqTnMVC6">/files/DpTjun7gVZipSqTnMVC6</a></td></tr><tr><td><strong>Skills</strong></td><td>Dynamically load prompts for the current goal.</td><td></td><td><a href="/pages/1y14YTSyAhbJKWugLqcw">/pages/1y14YTSyAhbJKWugLqcw</a></td><td><a href="/files/i9dLTP89QBsUoc0d1bUJ">/files/i9dLTP89QBsUoc0d1bUJ</a></td></tr><tr><td><strong>Python Library Setup</strong></td><td>Install and use our python library.</td><td></td><td><a href="/pages/2MUphEoXwKJgVQbP1CqJ">/pages/2MUphEoXwKJgVQbP1CqJ</a></td><td><a href="/files/SNr8RYg5pRIKq6dZeo9I">/files/SNr8RYg5pRIKq6dZeo9I</a></td></tr><tr><td><strong>Demo Project</strong></td><td>End to end demo video showing how to use Kiln</td><td></td><td><a href="/pages/qXaTfO53meIdxxsvPawY">/pages/qXaTfO53meIdxxsvPawY</a></td><td><a href="/files/KYOPmVKeWWKDnKRTGuOq">/files/KYOPmVKeWWKDnKRTGuOq</a></td></tr><tr><td><strong>Build a Reasoning Model</strong></td><td>Create your own reasoning model, like o3 or R1</td><td></td><td><a href="/pages/RaqHbAJUWPl2zk68NF3E">/pages/RaqHbAJUWPl2zk68NF3E</a></td><td><a href="/files/Q49HXGUBF6ORGBiL01mv">/files/Q49HXGUBF6ORGBiL01mv</a></td></tr></tbody></table>


# Quickstart

Download our app for MacOS, Windows or Linux

<figure><img src="/files/243fWEjvHpEs8TFLDfKO" alt=""><figcaption></figcaption></figure>

* :desktop: **macOS, Windows and Linux**: run our app anywhere.
* 🧑‍💻 **Intuitive UI**: Our UI is designed for anyone, from AI novices to experts.
* 🚀 **One Click Setup**: No docker, terminal or dependencies required.

### Step 1: Download and Install

MacOS and Windows: download the [latest release](https://github.com/Kiln-AI/Kiln/releases/latest). Be sure to download the correct version (Windows, Mac for Apple Silicon, Mac for Intel).

[![Download button](https://github.com/user-attachments/assets/a5d51b8b-b30a-4a16-a902-ab6ef1d58dc0)](https://github.com/Kiln-AI/Kiln/releases/latest)

Next, install the app:

* macOS: open the .dmg file, and drag the app to your Applications directory
* Windows: double click the installer, and follow the guide to install
* Linux: we suggest the installation command below.

<details>

<summary>Linux Installation Command</summary>

Run the following command in terminal to install Kiln:

{% code overflow="wrap" %}

```bash
curl -fsSL https://raw.githubusercontent.com/Kiln-AI/Kiln/refs/heads/main/app/desktop/linux_installer.sh | bash
```

{% endcode %}

You can download and review the install script before running. It will download the latest release from Github releases, install it to your path, and create an entry in your app launcher with a Kiln icon.

</details>

<details>

<summary>Windows Installer Troubleshooting</summary>

If you have any issues installing on Windows, check the solutions below:

* "Unrecognized App" warning: this warning may appear after we ship a new app update. This is just a warning. To bypass it, click "More Info" and continue the install.
* Anti-virus blockers: every release is scanned for viruses using [VirusTotal](https://www.virustotal.com/gui/home/upload), a comprehensive scanner that scans the app with over 70 virus scanners. Occasionally McAfee/AVG/Avast has a false positive (detects a virus when there isn't one). Feel free to scan the installer yourself (using VirusTotal or similar tool); once you trust the installer, add it to your virus scanner's allow-list.

</details>

### Step 2: Launch the App & Try Sample Task

Launch the app and get started!

The app will guide you through creating a project, creating a task, and connecting to AI providers like Ollama and OpenAI.

Try our sample task for a quick exploration, or define your own custom task for your project.

### Next: Video Guides and Docs

Our end to end demo walks you through all the major features of Kiln in a 20 minute video:

<p align="center"><a href="/pages/qXaTfO53meIdxxsvPawY"><img src="/files/JVUrr0EXdv6wSq79i0Xg" alt="Download button"></a></p>

Our documents explain each feature, and how to use it:

* [Evaluations](/docs/evals-and-specs/evaluations)
* [Synthetic Data Generation](/docs/synthetic-data-generation)
* [Fine Tuning Guide](/docs/fine-tuning/fine-tuning-guide)
* [Collaboration](/docs/collaboration)
* [Reviewing and Rating](/docs/reviewing-and-rating)
* [Organizing Datasets](/docs/organizing-datasets)
* See all docs in the navigation bar


# Models and AI Providers

Connect to Ollama, OpenAI, OpenRouter, AWS, Azure, Vertex, Together, Fireworks, and more

<figure><img src="/files/UJpczzzMbOD88QyF4pik" alt=""><figcaption></figcaption></figure>

## Skip the Guesswork: Kiln Suggests Models

Picking a model for a specific task can be hard. Each has different capabilities such as supported JSON modes, reasoning support, API features (logprobs, temperature) and censorship levels. Some work great for synthetic data and evals, others no so much.

We have written over 2000 test cases testing each popular model, on each AI provider, for each important feature. The tests are updated weekly, and we publish these capabilities in our model library. With this knowledge, the Kiln app automatically suggests appropriate models and model settings for each task. It will even warn you if you select a model that's unlikely to work. Read more on our blog: [I wrote 2000 LLM test cases so you don't have to](https://kiln.tech/blog/i_wrote_2000_llm_test_cases_so_you_dont_have_to).

<figure><img src="/files/l4dFIC3AyBEYPaVJphij" alt="" width="375"><figcaption><p>Kiln's model selector is task aware</p></figcaption></figure>

## Connecting AI Providers

When you first run Kiln, the app will prompt you to setup one or more AI providers. You need at least one for the core features of Kiln to function.

We currently support the following AI providers:

* Ollama
* OpenRouter
* OpenAI
* Groq
* Fireworks.ai
* Cerbras
* Together.ai
* AWS Bedrock
* Anthropic
* Gemini AI Studio / Gemini API
* Google Vertex AI
* Azure OpenAI
* HuggingFace
* Featherless AI
* SiliconFlow\.cn (for users in China)
* Any OpenAI compatible API like LiteLLM, vLLM, LMStudio, llama.cpp server, and many more

If you want to add or remove providers after initial setup, open `Settings > AI Providers & Models`.

{% hint style="success" %}
Don't see your provider listed? Most providers offer an OpenAI compatible API that Kiln can connect to. This includes common open source projects like LiteLLM and vLLM. Search their docs, and connect via the "Custom API" option in Kiln.
{% endhint %}

## Understanding and Adding Models

Models come in several flavours, from very easy to use, to advanced methods for expert users:

* [Included Models - Recommended](#included-models-recommended)
* [Custom Ollama models](#custom-ollama-models)
* [Custom models from existing providers](#custom-models-from-existing-providers)
* Provider Specific Guidance
  * [Azure OpenAI API](#azure-openai-api)
  * [Azure AI Foundry / Azure AI Studio](#azure-ai-foundry-formerly-azure-ai-studio-microsoft-ai-for-enterprise-360-elite)
* [Custom OpenAI compatible servers](#custom-openai-compatible-servers)
  * [LiteLLM](#litellm) - Anthropic, Huggingface, VertexAI, TogetherAI, and more.

### Included Models from the Model Library - Recommended

Included models are models that have been tested to work with Kiln's various features. These are the easiest to use, and generally won't result in errors.

To use these models simply connect any AI provider from the Settings page. Once connected, you can select these model from the model dropdown on the Run screen. The dropdown will warn if you attempt to use a model that doesn't support a feature (like structured output or synthetic data generation).

View all available models in our [model library on our webpage](https://kiln.tech/model_library) or the models tab in app. We update this list as new models come out.

<figure><img src="/files/0yilg0M5T2QpWgfUJtVN" alt="Kiln Model Library" width="350"><figcaption><p>The Kiln Model Library</p></figcaption></figure>

You can request we add models on our [Discord](https://kiln.tech/discord).

### Fine-Tuneable Models

The [model library](https://kiln.tech/model_library) or the models tab in app lists many of the of models that Kiln can fine-tune. Kiln can fine-tune even more models than shown in our library, including:

* Additional Fireworks.ai models: as soon as you connect a Fireworks.ai API key, over 60 additional models will be available for tuning in the Kiln UI. These are loaded live from Fireworks, and update automatically as new models are released. See a preview list below.
* Tune almost any model via tools like Unsloth: see our [fine-tuning guide](/docs/fine-tuning/fine-tuning-guide) for instructions on how to export fine-tuning datasets from Kiln for use with any tuning tool.

<details>

<summary>60+ Fine-Tuneable Models on Fireworks.ai</summary>

New models will automatically appear in Kiln as they are released by Fireworks. Here's a snapshot:

* Chronos Hermes 13B v2 (chronos-hermes-13b-v2)
* Code Llama 13B (code-llama-13b)
* Code Llama 13B Instruct (code-llama-13b-instruct)
* Code Llama 13B Python (code-llama-13b-python)
* Code Llama 34B (code-llama-34b)
* Code Llama 34B Instruct (code-llama-34b-instruct)
* Code Llama 34B Python (code-llama-34b-python)
* Code Llama 70B (code-llama-70b)
* Code Llama 70B Instruct (code-llama-70b-instruct)
* Code Llama 70B Python (code-llama-70b-python)
* Code Llama 7B (code-llama-7b)
* Code Llama 7B Instruct (code-llama-7b-instruct)
* Code Llama 7B Python (code-llama-7b-python)
* CodeQwen 1.5 7B (code-qwen-1p5-7b)
* Cogito v1 Preview Llama 3B (cogito-v1-preview-llama-3b)
* Cogito v1 Preview Llama 70B (cogito-v1-preview-llama-70b)
* Cogito v1 Preview Llama 8B (cogito-v1-preview-llama-8b)
* Cogito v1 Preview Qwen 14B (cogito-v1-preview-qwen-14b)
* Cogito v1 Preview Qwen 32B (cogito-v1-preview-qwen-32b)
* DeepSeek Coder 1.3B Base (deepseek-coder-1b-base)
* DeepSeek Coder 33B Instruct (deepseek-coder-33b-instruct)
* DeepSeek Coder 7B Base (deepseek-coder-7b-base)
* DeepSeek Coder 7B Base v1.5 (deepseek-coder-7b-base-v1p5)
* DeepSeek Coder 7B Instruct v1.5 (deepseek-coder-7b-instruct-v1p5)
* DeepSeek Coder V2 Instruct (deepseek-coder-v2-instruct)
* DeepSeek Coder V2 Lite Base (deepseek-coder-v2-lite-base)
* DeepSeek Coder V2 Lite Instruct (deepseek-coder-v2-lite-instruct)
* DeepSeek Prover V2 (deepseek-prover-v2)
* DeepSeek R1 (Fast) (deepseek-r1)
* Deepseek R1 05/28 (deepseek-r1-0528)
* DeepSeek R1 0528 Distill Qwen3 8B (deepseek-r1-0528-distill-qwen3-8b)
* DeepSeek R1 (Basic) (deepseek-r1-basic)
* DeepSeek R1 Distill Llama 70B (deepseek-r1-distill-llama-70b)
* DeepSeek R1 Distill Llama 8B (deepseek-r1-distill-llama-8b)
* DeepSeek R1 Distill Qwen 14B (deepseek-r1-distill-qwen-14b)
* DeepSeek R1 Distill Qwen 1.5B (deepseek-r1-distill-qwen-1p5b)
* DeepSeek R1 Distill Qwen 32B (deepseek-r1-distill-qwen-32b)
* DeepSeek R1 Distill Qwen 7B (deepseek-r1-distill-qwen-7b)
* DeepSeek V2 Lite Chat (deepseek-v2-lite-chat)
* DeepSeek V2.5 (deepseek-v2p5)
* DeepSeek V3 (deepseek-v3)
* Deepseek V3 03-24 (deepseek-v3-0324)
* Dolphin 2.9.2 Qwen2 72B (dolphin-2-9-2-qwen2-72b)
* FireFunction V2 (firefunction-v2)
* Llama 4 Maverick Instruct (Basic) (llama4-maverick-instruct-basic)
* Llama 4 Scout Instruct (Basic) (llama4-scout-instruct-basic)
* Llama Guard v2 8B (llama-guard-2-8b)
* Llama Guard v3 1B (llama-guard-3-1b)
* Llama Guard 3 8B (llama-guard-3-8b)
* Llama Guard 7B (llamaguard-7b)
* Llama 2 13B (llama-v2-13b)
* Llama 2 13B Chat (llama-v2-13b-chat)
* Llama 2 70B (llama-v2-70b)
* Llama 2 70B Chat (llama-v2-70b-chat)
* Llama 2 7B (llama-v2-7b)
* Llama 2 7B Chat (llama-v2-7b-chat)
* Llama 3 70B Instruct (llama-v3-70b-instruct)
* Llama 3 70B Instruct (HF version) (llama-v3-70b-instruct-hf)
* Llama 3 8B (llama-v3-8b)
* Llama 3 8B Instruct (llama-v3-8b-instruct)
* Llama 3 8B Instruct (HF version) (llama-v3-8b-instruct-hf)
* Llama 3.1 70B Instruct (llama-v3p1-70b-instruct)
* Llama 3.1 8B Instruct (llama-v3p1-8b-instruct)
* Llama 3.1 Nemotron 70B (llama-v3p1-nemotron-70b-instruct)
* Llama 3.2 1B (llama-v3p2-1b)
* Llama 3.2 1B Instruct (llama-v3p2-1b-instruct)
* Llama 3.2 3B (llama-v3p2-3b)
* Llama 3.2 3B Instruct (llama-v3p2-3b-instruct)
* Llama 3.3 70B Instruct (llama-v3p3-70b-instruct)
* MythoMax L2 13B (mythomax-l2-13b)
* Nous Hermes 2 Yi 34B (nous-hermes-2-yi-34b)
* Nous Hermes Llama2 13B (nous-hermes-llama2-13b)
* Nous Hermes Llama2 70B (nous-hermes-llama2-70b)
* Nous Hermes Llama2 7B (nous-hermes-llama2-7b)
* Phind CodeLlama 34B Python v1 (phind-code-llama-34b-python-v1)
* Phind CodeLlama 34B v1 (phind-code-llama-34b-v1)
* Phind CodeLlama 34B v2 (phind-code-llama-34b-v2)
* Qwen1.5 72B Chat (qwen1p5-72b-chat)
* Qwen2 72B Instruct (qwen2-72b-instruct)
* Qwen2 7B Instruct (qwen2-7b-instruct)
* Qwen2.5 0.5B Instruct (qwen2p5-0p5b-instruct)
* Qwen2.5 14B (qwen2p5-14b)
* Qwen2.5 14B Instruct (qwen2p5-14b-instruct)
* Qwen2.5 1.5B Instruct (qwen2p5-1p5b-instruct)
* Qwen2.5 32B (qwen2p5-32b)
* Qwen2.5 32B Instruct (qwen2p5-32b-instruct)
* Qwen2.5 72B (qwen2p5-72b)
* Qwen2.5 72B Instruct (qwen2p5-72b-instruct)
* Qwen2.5 7B (qwen2p5-7b)
* Qwen2.5 7B Instruct (qwen2p5-7b-instruct)
* Qwen2.5-Coder 0.5B (qwen2p5-coder-0p5b)
* Qwen2.5-Coder 0.5B Instruct (qwen2p5-coder-0p5b-instruct)
* Qwen2.5-Coder 14B (qwen2p5-coder-14b)
* Qwen2.5-Coder 14B Instruct (qwen2p5-coder-14b-instruct)
* Qwen2.5-Coder 1.5B (qwen2p5-coder-1p5b)
* Qwen2.5-Coder 1.5B Instruct (qwen2p5-coder-1p5b-instruct)
* Qwen2.5-Coder 32B (qwen2p5-coder-32b)
* Qwen2.5-Coder 32B Instruct (qwen2p5-coder-32b-instruct)
* Qwen2.5-Coder 32B Instruct 128K (qwen2p5-coder-32b-instruct-128k)
* Qwen2.5-Coder 32B Instruct 32K RoPE (qwen2p5-coder-32b-instruct-32k-rope)
* Qwen2.5-Coder 32B Instruct 64k (qwen2p5-coder-32b-instruct-64k)
* Qwen2.5-Coder 3B (qwen2p5-coder-3b)
* Qwen2.5-Coder 3B Instruct (qwen2p5-coder-3b-instruct)
* Qwen2.5-Coder 7B (qwen2p5-coder-7b)
* Qwen2.5-Coder 7B Instruct (qwen2p5-coder-7b-instruct)
* Qwen2.5-Math 72B Instruct (qwen2p5-math-72b-instruct)
* Qwen2.5-VL 32B Instruct (qwen2p5-vl-32b-instruct)
* Qwen2.5-VL 3B Instruct (qwen2p5-vl-3b-instruct)
* Qwen2.5-VL 72B Instruct (qwen2p5-vl-72b-instruct)
* Qwen2.5-VL 7B Instruct (qwen2p5-vl-7b-instruct)
* Qwen3 0.6B (qwen3-0p6b)
* Qwen3 1.7B (qwen3-1p7b)
* Qwen3 32B (qwen3-32b)
* Qwen3 4B (qwen3-4b)
* Qwen3 8B (qwen3-8b)
* Qwen QWQ 32B Preview (qwen-qwq-32b-preview)
* Qwen2.5 14B Instruct (qwen-v2p5-14b-instruct)
* Qwen2.5 7B (qwen-v2p5-7b)
* QWQ 32B (qwq-32b)
* Rolm OCR (rolm-ocr)
* Yi 34B (yi-34b)
* Nous Capybara 34B V1.9 (yi-34b-200k-capybara)
* Yi 34B Chat (yi-34b-chat)

</details>

### Custom Ollama Models

Any Ollama model you have installed on your server will be available to use in Kiln. To add models, simply install them with the Ollama CLI `ollama pull <model_name>`.

Some Ollama models are included/tested, and will automatically appear in the model dropdown. Any untested Ollama models will still appear in the dropdown, but in the "Untested" section.

### Custom Models from Existing Providers

If you want to use a model that is not in the list but is supported by one of our AI providers, you can use a custom model.

To use a custom model, click "Add Model" in the "AI Providers & Models" section of Settings.

These will appear in the "untested" section of the model dropdown.

### Provider Specific Guidance

#### Azure OpenAI API

When using Azure OpenAI API, you need to deploy each model you want to use, manually through the Azure console. If you have not, you'll get deployment errors when trying to call a model.

* **Suggested - Deploy with Default Names**: If you deploy with the default names, for example "gpt-4o"/"gpt-4o-mini", you can simply use the models using the dropdown in Kiln.
* **Deployments with Custom Names**: If you have a non-standard deployment name, you'll have to add each model as a [custom model](#custom-models-from-existing-providers), using the deployment name as the model name.

#### Azure AI Foundry (formerly Azure AI Studio, Microsoft AI for Enterprise 360 Elite)

When using Azure AI Foundry, you need to deploy each model you want to use manually through the Azure console. If you have not, you'll get deployment errors when trying to call a model.

After deploying a model, you must add it to Kiln as a [custom model](#custom-models-from-existing-providers), using the deployment name as the model name.

#### Google Vertex AI

When using Vertex, many models need to be manually enabled through the console before using them (primarily Anthropic models). If you see errors when trying to run a model, open the vertex AI console for your project, go to the model garden, and enable that model.

Similarly, if you see quota errors you may need to manage/request quota from the Vertex console. Quota is specific to the model + region. Ensure you request quota in the region you specified when you connected Vertex AI to Kiln.

#### Hugging Face

Hugging face has thousands of models. We've included a few of these common models in the Kiln built-in model list, but you can add any hugging face model via the [custom model](#custom-models-from-existing-providers) option.

Hugging face errors are not always descriptive - if you get 400 errors, it's likely the model you've selected requires a Hugging Face Pro subscription. Try the same model in their UI for a more helpful error message.

### Custom OpenAI Compatible Servers

If you have an OpenAI compatible server (LiteLLM, vLLM, etc.), you can use it in Kiln.

To do this, add a "Custom API" in the "AI Providers & Models" section of Settings.

All models supported by this API will appear in the "untested" section of the model dropdown.

Notes:

* The API must support the `/v1/models` endpoint, so Kiln can access the list of models.
* Many Kiln tasks produce structured (JSON) output. These can be hard to get working on custom servers, as each server/model pair usually needs some configuration to reliably produce structured output (tools vs json\_mode vs json parsing vs json\_schema, etc).


# End to End Project Demo

Video walkthrough of creating tasks, evals, fine-tuning, synthetic data gen, and more!

The demo video below walks you through every step of creating a project in Kiln. It's a great place to start if you prefer video walk throughs.

{% hint style="info" %}
See the relevant docs pages for more comprehensive documentation and demo videos of each feature. This video is a quick high level walkthrough.
{% endhint %}

The video covers:

* [Creating an eval](/docs/evals-and-specs/evaluations) including generating synthetic eval data, creating LLM-as-judge evals, and validating the eval with human ratings
* [Finding the best way to run your task](/docs/evals-and-specs/evaluations#finding-the-ideal-run-method) by evaluating prompt/model pairs
* [Fine-tuning models](/docs/fine-tuning/fine-tuning-guide) including synthetic training data and evaluating tunes
* [Iterating as project evolves](https://kiln.tech/blog/you_need_many_small_evals_for_ai_products#setup-team-evals) new evals and prompts
* [Set up collaboration with git & GitHub](/docs/collaboration)

{% embed url="<https://vimeo.com/1104945621>" %}


# Optimizers

Tools to build the optimal AI system for your task

Kiln offers a powerful set of optimization techniques to make your AI agents smarter, faster, and cheaper.

Click the "Optimize" tab in the app to explore our optimizers, or browse the docs:

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Optimize Prompt</strong></td><td>Improve your prompt using our automatic optimizers or manually.</td><td><a href="/files/CYsusrOr631L4LlBM60z">/files/CYsusrOr631L4LlBM60z</a></td><td><a href="/pages/ZZcdUMMQOzSwNwVTPdqF">/pages/ZZcdUMMQOzSwNwVTPdqF</a></td></tr><tr><td><strong>Compare Models</strong></td><td>Compare models to find the best quality/cost tradeoff for your task.</td><td><a href="/files/kJti9HACvdgX2vjtVpml">/files/kJti9HACvdgX2vjtVpml</a></td><td><a href="https://kiln.tech/model_library">https://kiln.tech/model_library</a></td></tr><tr><td><strong>Fine-Tuning</strong></td><td>Train a custom model for your task.</td><td><a href="/files/reaAySjXgTjdVkcX2Tzy">/files/reaAySjXgTjdVkcX2Tzy</a></td><td><a href="/pages/d10ZmwweJ8LYSmlXiLB6">/pages/d10ZmwweJ8LYSmlXiLB6</a></td></tr><tr><td><strong>Docs &#x26; Search (RAG)</strong></td><td>Let agents search for relevant knowledge.</td><td><a href="/files/opZURxUeL4URf4LA99i5">/files/opZURxUeL4URf4LA99i5</a></td><td><a href="/pages/Dtsf7EZfJgPB9VbSUiae">/pages/Dtsf7EZfJgPB9VbSUiae</a></td></tr><tr><td><strong>Tools &#x26; MCP</strong></td><td>Add tools like web-search, code sandboxes, and more.</td><td><a href="/files/DrxYj7KvQVl0GRHKPQea">/files/DrxYj7KvQVl0GRHKPQea</a></td><td><a href="/pages/qmzsaLmyoO8DcirfJpAK">/pages/qmzsaLmyoO8DcirfJpAK</a></td></tr><tr><td><strong>Agents</strong></td><td>Allow your task to call sub-agents and perform work.</td><td><a href="/files/HUfoCrZmOkhj4AOFybMJ">/files/HUfoCrZmOkhj4AOFybMJ</a></td><td><a href="/pages/6X2jCKYVfnoJLkVMp41g">/pages/6X2jCKYVfnoJLkVMp41g</a></td></tr><tr><td><strong>Skills</strong></td><td>Dynamically load prompts for the current goal.</td><td><a href="/files/i9dLTP89QBsUoc0d1bUJ">/files/i9dLTP89QBsUoc0d1bUJ</a></td><td></td></tr></tbody></table>


# Fine Tuning

Create custom fine-tuned models for your use case

Kiln makes it easy to fine-tune a wide variety of models like GPT-4o, Llama, Mistral, Gemma, and many more. Check out our fine tuning docs to get started:

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Fine Tuning Guide</strong></td><td>Our end-to-end walkthrough of fine-tuning a model in Kiln. Includes generating training data.</td><td><a href="/files/reaAySjXgTjdVkcX2Tzy">/files/reaAySjXgTjdVkcX2Tzy</a></td><td><a href="/pages/UQ5mUIArq32hihhGE4Sv">/pages/UQ5mUIArq32hihhGE4Sv</a></td></tr><tr><td><strong>Train a Reasoning Model</strong></td><td>Using a process called distillation, create your own reasoning model.</td><td><a href="/files/Q49HXGUBF6ORGBiL01mv">/files/Q49HXGUBF6ORGBiL01mv</a></td><td><a href="/pages/RaqHbAJUWPl2zk68NF3E">/pages/RaqHbAJUWPl2zk68NF3E</a></td></tr><tr><td><strong>Fine Tune For Tool Use</strong></td><td>Train a model to use a specific set of tools, at the right time, with the right parameters.</td><td><a href="/files/RqmwElvFlPnm4q3vJ1Nq">/files/RqmwElvFlPnm4q3vJ1Nq</a></td><td><a href="/pages/SbUpu30PLNTEjXCKTNOK">/pages/SbUpu30PLNTEjXCKTNOK</a></td></tr></tbody></table>

{% hint style="success" %}
**Before Fine-Tuning, consider using Kiln's** [**Automatic Prompt Optimizer**](/docs/prompts/automatic-prompt-optimizer)**.** It has many of the benefits of fine-tuning (automated data driven optimization), while the result is much easier to deploy than a fine-tuned model.
{% endhint %}


# Fine Tuning Guide

Fine tuning 9 Models in 18 minutes

Kiln makes it easy to fine-tune a wide variety of models like GPT-4o, Llama, Mistral, Gemma, and many more.

### Overview

In this guide we'll be walking through an example where we start from scratch, and build 9 fine-tuned models in just under 18 minutes of active work, not counting time waiting for training and data-gen.

No coding is necessary, and our UI will guide you through the process. Step 6 is optional, and requires some basic python skills. Our open-source python library is available for advanced users.

You can follow this guide to create your own LLM fine-tunes. We'll cover:

A Demo Project:

* \[2 mins]: [Define task, goals, and schema](#step-1-define-your-task-and-goals)
* \[9 mins]: [Synthetic data generation](/docs/synthetic-data-generation): create 920 high-quality examples for training
* \[5 mins]: Dispatch 9 fine tuning jobs: [Fireworks](#step-4-dispatch-training-jobs) (Llama 3.2 1b/3b/11b, Llama 3.1 8b/70b, Mixtral 8x7b), [OpenAI](#step-4-dispatch-training-jobs) (GPT 4o, 4o-Mini), and [Unsloth](#step-6-optional-training-on-your-own-infrastructure) (Llama 3.2 1b/3b). Note: since this guide was written we've added over 60 new models for fine tuning!
* \[2 mins]: [Deploy your new models and test they work](#step-5-deploy-and-run-your-models)

Analysis:

* [Cost Breakdown](#cost-breakdown)
* [Next steps](#next-steps): evaluation, exporting models, iteration and data strategies

{% hint style="info" %}
If you want to tune a reasoning model, see our [guide for training reasoning models](/docs/fine-tuning/guide-train-a-reasoning-model). It includes notes about each step of this guide which are necessary to produce a reasoning model.
{% endhint %}

### Step 1: Define your Task

First, we’ll need to define what the models should do. In Kiln we call this a “task definition”. Create a new task in the Kiln UI to get started, including a initial prompt and input/output schema.

For this demo we'll make a task that generates news article headlines of various styles, given a summary of a news topic.

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FEJ4b8A4QiEQlOGbYkXDX%2Fuploads%2FL6pUgqBP3wgAhIZtfJ5L%2FCreateTask720.mp4?alt=media&token=1e8dcf3a-bb10-4774-a421-0fed5b0adb2e>" %}
Create a task to fine-tune for
{% endembed %}

### Step 2: Generate Training Data with Synthetic Data Generation

To fine tune, you’ll need a dataset to learn from.

Kiln offers an interactive UI for quickly and easily building synthetic datasets. In the video below we use it to generate 920 training examples in 9 minutes of hands-on work. See our [data gen guide](/docs/synthetic-data-generation) for more details.

{% hint style="info" %}
All fine tuning data must be tagged with a [tag](/docs/organizing-datasets#using-tags-to-organize-your-dataset) starting with `fine_tune` (e.g. fine\_tune, fune\_tune\_thinking, fine\_tune\_experiment\_42).

If you launch synthetic data gen from within the "Create a New Fine Tune" screen, the tag `fine_tune` will automatically be added to all generated samples.

If you already created tuning data, use the dataset tab to add the `fine_tune` tag to the samples you want to use for tuning.
{% endhint %}

Kiln includes topic trees to generate diverse content, a range of models/prompting strategies, interactive guidance and interactive UI for curation/correction.

When generating synthetic data you want to generate the best quality content possible. Don’t worry about cost and performance at this stage. Use large high quality models, detailed prompts with multi-shot prompting, chain of thought, and anything else that improves quality. You’ll be able to address performance and costs in later steps with fine tuning.

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FEJ4b8A4QiEQlOGbYkXDX%2Fuploads%2FwVRCYmuIg1s2d38NLZ2t%2FDatagen720.mp4?alt=media&token=6d1a18c7-67a7-4062-8ac6-fb88525676e7>" %}
Synthetic Data Generation
{% endembed %}

### Step 3: Select Models to Fine Tune

Kiln supports over 60 fine-tuneable models using three different service based tuning providers:

* Open AI: GPT 4.1, 4o, 4.1-mini and 4o-mini
* Google Gemini: Gemini 2.0 flash and Gemini 2.0 Pro
* Fireworks.ai: over 60 open weight models including Qwen 2.5, Llama 2/3.x, Deepseek V3/R1, QwQ, and more. See the [full list here](/docs/models-and-ai-providers#additional-fine-tuneable-models).
* Together AI: Llama 3.1 8b/70b, Llama 3.2 1b/3b, Qwen2.5 14b/72b

{% hint style="success" %}
To see more options on the "Create Fine Tune" screen, connect API keys for the providers listed above in Settings.
{% endhint %}

For this demo we choose 9 models to experiment with.

### Step 4: Dispatch Training Jobs

Use the "Fine Tune" tab in the Kiln UI to kick off your fine-tunes. Simply select the models you want to train, select a dataset, and add any training parameters.

{% hint style="info" %}
**Training Reasoning/Thinking Model**

Kiln can train a reasoning model. See the guide on [training reasoning models](/docs/fine-tuning/guide-train-a-reasoning-model).
{% endhint %}

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FEJ4b8A4QiEQlOGbYkXDX%2Fuploads%2F90AZYx6JVoCYFrhHCTUx%2FCreateTrainingJobs720.mp4?alt=media&token=c0e76a1b-0f55-4297-85b9-b1f9b6c4125f>" %}
Dispatching Training Jobs. Note: video does not match current UI
{% endembed %}

### Step 5: Deploy and Run Your Models

Kiln will automatically deploy your fine-tunes when they are complete. You can use them from the Kiln UI without any additional configuration. Simply select a fine-tune by name from the model dropdown in the "Run" tab.

Together, Fireworks and OpenAI tunes are deployed "serverless". You only pay for usage (tokens), with no recurring costs.

You can use your models outside of Kiln by calling Fireworks or OpenAI APIs with the model ID from the "Fine Tune" tab.

**Early Results**: Our fine-tuned models show some immediate promise. Previously models smaller than Llama 70b failed to produce the correct structured data for our task. After fine tuning even the smallest model, Llama 3.2 1b, consistently works.

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FEJ4b8A4QiEQlOGbYkXDX%2Fuploads%2FevKBtLibJBN56hVGULL7%2FRun720.mp4?alt=media&token=f072c028-088b-4673-9d10-f88580d71cb7>" %}
Running some of our Fine Tuned Models (Llama 3.2 1b & GPT 4o-mini)
{% endembed %}

{% hint style="info" %}
If a Fireworks fine tune gives you the error \`Model not found, inaccessible, and/or not deployed\`, it means that model was un-deployed by Fireworks. Opening the model in the "Fine Tune" tab of Kiln will trigger a re-deploy.
{% endhint %}

### Step 6 \[Optional]: Training on your own Infrastructure

Kiln can also export your dataset to common formats for fine tuning on your own infrastructure. Simply select one of the "Download" options when creating your fine tune, and use the exported JSONL file to train with your own tools.

We currently recommend [Unsloth](https://github.com/unslothai/unsloth) and Axolotl. These platforms let you train almost any open model, including Gemma, Mistral, Llama, Qwen, Smol, and [many more](https://docs.unsloth.ai/get-started/all-our-models).

**Unsloth Example**

See this example [Unsloth notebook](https://colab.research.google.com/drive/1Ivmt4rOnRxEAtu66yDs_sVZQSlvE8oqN?usp=sharing), which has been modified to load a dataset file exported from Kiln. You can use it to fine-tune locally or in Google Colab.

Export your dataset using the "Hugging Face chat template (JSONL)" option for compatibility with the demo notebook.

{% embed url="<https://files.gitbook.com/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FEJ4b8A4QiEQlOGbYkXDX%2Fuploads%2Fsr97NzR4hJGsUWiIJPCv%2FUnsloth720.mp4?alt=media&token=041646ce-fbeb-4627-894b-1b4a20f27090>" %}
Unsloth Demo
{% endembed %}

#### Google Gemini on Vertex AI

Kiln can generate the training format needed by Google's Vertex AI to fine tune Gemini models.

Select the Vertex AI/Gemini option in the dropdown, to download training/validation files in the appropriate format. Then follow the Google's fine-tuning [guide](https://cloud.google.com/vertex-ai/generative-ai/docs/models/tune-models), using the files from Kiln as your training/validation sets.

### Cost Breakdown

Our demo use case was quite reasonably priced.

| Task                                  | Platform                   | Cost (USD) |
| ------------------------------------- | -------------------------- | ---------- |
| Training Data Generation              | OpenRouter                 | $2.06      |
| Fine-tuning 5x Llama models + Mixtral | Fireworks                  | $1.47      |
| Fine-tuning GPT-4o Mini               | OpenAI                     | $2.03      |
| Fine-tuning GPT-4o                    | OpenAI                     | $16.91     |
| Fine-tuning Llama 3.2 (1b & 3b)       | Unsloth on Google Colab T4 | $0.00      |

If it wasn't for GPT-4o, the whole project would have cost less than $6!

Meanwhile our fastest fine-tune (Llama 3.2 1b) is about 10x faster and 150x cheaper than the models we used during synthetic data generation (source:OpenRouter perf stats & prices).

### Track Training Metrics with Weights & Biases

Kiln supports tracking training metrics with the tool [Weights & Biases](https://wandb.ai/site/) . Configure your W\&B API key in `Settings > AI Providers & Models > Weights & Biases` before starting your fine-tuning job. Metrics will appear for any training jobs on Fireworks or Together. OpenAI doesn't support W\&B, but provides similar metrics in their own dashboard, which is linked from the Kiln Fine Tune page.

<figure><img src="/files/VsScVkLMDfVKT34cG4Ca" alt="" width="287"><figcaption><p>Weights and Biases Metrics</p></figcaption></figure>

### Next Steps

What’s next after fine tuning?

#### **Evaluate Model Quality**

We now have 9 fine-tuned models, but which is best for our task? We should evaluate them for quality/speed/cost tradeoffs.

Kiln has [powerful evaluation tools](/docs/evals-and-specs/evaluations) to help you though this process. Check out the [evaluation guide](/docs/evals-and-specs/evaluations) for details.

If your task is deterministic (classification), Kiln AI will provide the validation set to OpenAI or Together during tuning, and they will report val\_loss on their dashboard. For non-deterministic tasks (including generative tasks) you can use our [evaluation tools](/docs/evals-and-specs/evaluations) to evaluate quality.

#### **Exporting Models**

You can export your models for use on your machine, deployment to the cloud, or embedding in your product.

* Fireworks: you can [download the weights](https://docs.fireworks.ai/fine-tuning/fine-tuning-models#downloading-model-weights) in Hugging Face PEFT format, and convert as needed.
* Together: you can [download the weights](https://docs.together.ai/docs/finetuning#running-your-model-locally), run locally or convert as needed.
* Unsloth: your fine-tunes can be directly exported to GGUF or other formats which make these model easy to deploy. A GGUF can be [imported to Ollama](https://github.com/ollama/ollama/blob/main/docs/import.md) for local use. Once added to Ollama, the models will become available in Kiln UI as well.
* OpenAI: sadly OpenAI won’t let you download their models.

#### **Iterate to Improve Quality**

Models and products are rarely perfect on their first try. When you find bugs or have new goals, Kiln makes it easy to build new models. Some ways to iterate:

* Experiment with different base-models
* Experiment with fine-tuning hyperparameters (see the "Advanced Options" section of the UI)
* Experiment with shorter training prompts, which can reduce costs
* For one-off bugs you encounter use Kiln to “[repair](/docs/repairing-responses)” the issues. These get added to your training data for future iterations.
* For recurring bugs/patterns, use [synthetic data generation](/docs/synthetic-data-generation) to generate many samples of common bugs, ensure they have correct responses with [human guidance](/docs/synthetic-data-generation#templates-and-custom-guidance), and add the results to the training set to prevent this class of issues in the future.
* Rate your dataset using Kiln’s [rating system](/docs/reviewing-and-rating), then build fine-tunes using only highly rated content.
* Regenerate fine-tunes as your dataset grows and evolves
* Try new foundation models (directly and with fine tuning) when new state of the art models are released.

#### **Integrate with Code**

Kiln can be used entirely through the UI and doesn't require coding. However, if you'd like a code-based integration, our open-source [Python library](https://pypi.org/project/kiln-ai/) is available. You can dispatch fine-tune jobs and call fine-tuned models through our library without the UI if you prefer.

### **Our "Ladder" Data Strategy**

Kiln enables a "Ladder" data strategy: the steps start from small quantity and high effort, and progress to high quantity and low effort. Each step builds on the prior:

* \~10 manual high quality examples.
* \~30 LLM generated examples using the prior examples for multi-shot prompting. Use expensive models, detailed prompts, and token-heavy techniques (chain of thought). Manually review each ensuring low quality examples are not used as samples.
* \~1000 synthetically generated examples, using the prior content for multi-shot prompting. Again, using expensive models, detailed prompts and chain of thought. Some interactive sanity checking as we go, but less manual review once we have confidence in the prompt and quality.
* 1M+: after fine-tuning on our 1000 sample set, most inference happens on our fine-tuned model. This model is faster and cheaper than the models we used for building it through zero shot prompting, shorter prompts, and smaller models.

Like a ladder, skipping a step is dangerous. You need to make sure you’re solid before you continue to the next step.


# Train a Reasoning Model

Built a reasoning model, like o3 or R1, for your use case

Kiln is a platform that makes building task-specific AI models easy and fast. By creating a fine-tuned model targeted to your use case, you can produce a model that's higher quality, faster and cheaper than standard foundation models.

In this guide, we'll walk through how to build a reasoning model, like OpenAI o3 or Deepseek R1, for your specific use case. The whole process can be completed in as little as 30 minutes, and does not require coding.

Also see our docs on [how Kiln supports Reasoning and Chain of Thought](/docs/reasoning-and-chain-of-thought).

## How to Train a Task Specific Reasoning Model

We already have a [detailed guide on fine-tuning models](/docs/fine-tuning/fine-tuning-guide). This article covers the settings to use throughout that process to ensure the final model you produce is a reasoning model, and not just a standard fine-tune.

### Training Demo Video

{% embed url="<https://vimeo.com/1067942685/f2fbe3e3de>" %}
Video walkthrough of creating a custom reasoning LLM
{% endembed %}

### Ensure your training data includes "reasoning"

When developing your training data with our [synthetic data generation](/docs/synthetic-data-generation) tool, be sure to use either a reasoning model or chain-of-thought prompting. Using either of these will ensure your dataset has reasoning data to learn from. See our [model list](/docs/models-and-ai-providers#included-models-recommended) for which models have native reasoning support.

See below for [how to choose between reasoning and chain of thought](#choosing-between-reasoning-and-chain-of-thought).

<figure><img src="/files/ZjyZLfrMkrKamvMQLWG6" alt="" width="375"><figcaption><p>Synthetic Data Run Options</p></figcaption></figure>

{% hint style="info" %}
If you're using multi-shot prompting, also ensure your prompt examples have appropriate reasoning data. Consider a [custom prompt](/docs/prompts#custom-prompts-saved-prompts) with examples demonstrating ideal reasoning for your task.
{% endhint %}

### Create a training dataset filtered to samples with reasoning

When creating your fine-tuning dataset, be sure to filter it to samples with reasoning/thinking checking "Filter to Reasoning Samples" as shown here:

<figure><img src="/files/CaT0I57nK6EtLXbGMNy9" alt="" width="375"><figcaption><p>Filtering your training dataset</p></figcaption></figure>

### Choose the correct training strategy

To train your own reasoning model, you must select the `Thinking - Learn both thinking and final response` in the `Reasoning` dropdown of Step 3. This will include the reasoning data when fine-tuning.

<figure><img src="/files/Lw5Hlusf84DnosSwPg2d" alt="" width="375"><figcaption><p>Select this on the "Create Fine Tune" screen</p></figcaption></figure>

If you select `Disabled` the fine-tune will only learn from the final result, not the reasoning. This is still a valid approach and could produce a viable model for your task. However, it won't produce a model with learned reasoning skills.

### Call your fine-tuned model with the appropriate prompt

When you call any fine-tune, we always recommend calling it with the same prompt used in training.

<figure><img src="/files/rDOjiEP9hJ1BjaqeDD2w" alt="" width="341"><figcaption><p>Kiln's Inference UI Options</p></figcaption></figure>

If calling your model from custom code, follow the chat call flow described in our [reasoning docs](/docs/reasoning-and-chain-of-thought#chain-of-thought-call-flow-non-reasoning-model), and use the prompts used to fine-tune the model which can be found by clicking the model in the "Fine Tune" tab of Kiln's UI.

{% hint style="info" %}
If you want to call a fine-tune with a shorter prompt for performance reasons, consider training it with that prompt — even if sample data was generated with a longer prompt. The results will typically be better than using a prompt the model didn't see at training time.

To do this select "Custom Prompt" when creating the fine-tune, and set your prompt there.
{% endhint %}

### Choosing between Reasoning and Chain of Thought

The fine-tuning approach described in this article is a general approach that trains for an intermediate "thinking" output. This can be used for both reasoning models and chain of thought.

{% hint style="info" %}
Read about [the difference between reasoning models and chain of thought](/docs/reasoning-and-chain-of-thought#what-are-reasoning-models-and-chain-of-thought).
{% endhint %}

Both approaches can build great task specific models. Which to choose depends on your use case. It can be worth training and comparing several to find the best option for you.

* **Distill a Reasoning Model**: Reasoning models have learned reasoning skills across a range of domains. If large reasoning models like Deepseek R1 perform well on your task, but are too expensive or slow, it can be a good choice to fine-tune a smaller model from R1 outputs (this is called distilling a model). The smaller model will learn task-specific reasoning patterns from R1 samples, and be faster and cheaper to run.
* **Chain of thought with default prompt**: Sometimes a simple "think step by step" prompt is all you need for chain of thought to greatly improve your quality of output. If large models work great with a simple prompt but smaller models fail to produce the same quality, you can build a fine-tune with task-specific examples so the smaller model can distill the thinking patterns from the larger model.
* **Chain of thought with a custom thinking prompt**: When building a model for a specific task, it's very possible you or your team understand the nuance of the task better than a generalized model like Deepseek R1. If you can create a "thinking instructions" prompt that works well with large models like Sonnet or GPT-4o, you can use that to build a synthetic training set, create a fine-tune, and reproduce that quality on a much smaller and faster model.

In each case, you're building a model that will be focused on the use-case samples it is trained on. This can produce a model that's faster, cheaper and higher quality than the original model, within the domain of your task.

{% hint style="info" %}
For the example in the demo video, I actually found custom chain-of-thought on Sonnet 3.5 to produce better content than Deepseek R1. It's worth experimenting with different models and prompts to find the best pair suited for your task, before jumping to building a synthetic dataset.
{% endhint %}

### Improving quality with human curation and feedback

Human curation feedback can add the nuance that makes a truly great model/product. Kiln offers a number of tools to make this easy:

* Have a subject matter expert [rate the synthetic training set](/docs/reviewing-and-rating), and filter your training data to only use high quality samples.
* Have subject matter experts [repair poorly rated outputs](/docs/repairing-responses), giving the model important examples of places it likely would have failed without fine-tuning.
* Use human-led chain of thought prompts as described [here](#choosing-between-reasoning-and-chain-of-thought), to generate [large synthetic data sets](/docs/synthetic-data-generation) for fine-tuning.
* When you find a pattern of bugs, use [synthetic data generation with human guidance](/docs/synthetic-data-generation) to create samples of correct input/output pairs. Add these to your training set to fix the behaviour the next time you train.
* Use Kiln's [collaboration system](/docs/collaboration) to allow anyone on your team to contribute to model quality with feedback, data generation and quality. Our UI is designed for anyone, and does not require command line or coding skills.


# Fine Tuning for Tool Use

Build fine-tuned models for calling a set of tools, like MCP

Kiln can fine-tune a model for calling a specific set of tools. The fine-tuned models can improve over the base model by:

* Learning when to call each tool, and when not to
* Learning when to choose one tool over another
* Learning how to format tool calls and tool call parameters. This greatly reduces error rates over the base model, making smaller and faster models viable.

Together, this means you can improve agent performance and lower costs.

#### Building a Fine-Tuning Training Dataset for Tool Calling

To create a fine-tune targeting tool calling, you must generate a training set specifically for tool calling.

{% hint style="info" %}
The tool set available during training data generation must exactly match the tool set your fine-tune targets.

Kiln disallows training on samples that don't have a matching toolset. We don't want to train on these as the fine-tuned model would improperly learn not to call a tool, even when a tool call would have been appropriate.

This doesn’t mean every tool needs to be called in every training sample. Only that every tool was available to be called.
{% endhint %}

Kiln makes building a tool-calling training dataset easy:

1. Open the Fine-Tune tab.
2. Click “Create Fine Tune”.
3. Select the set of tools that the model should learn to call.

   <figure><img src="/files/QNt1cbI2e5z4xnlUwL3e" alt="" width="375"><figcaption><p>Selecting tools available to the fine-tuned model</p></figcaption></figure>
4. Click “Add Fine-Tuning Data” to launch Kiln's synthetic data generation tool.

   <figure><img src="/files/WbLiVGkMsEqq1YFgDxrS" alt="" width="375"><figcaption></figcaption></figure>
5. Generate synthetic training data using Kiln's [synthetic data gen](/docs/synthetic-data-generation) tool. It will automatically select the correct tools for you when generating sample outputs.

#### Distilling Larger Models and Longer Prompts for Better Tool Calling

Beyond learning tool-call formatting, your fine-tuned model must learn *when* to call each tool and *which parameters* to pass. The quality of the resulting model depends heavily on the quality of the training dataset—so how do you build a high-quality dataset for tool calling?

The answer is typically [*distillation*](https://en.wikipedia.org/wiki/Knowledge_distillation): training a model on the outputs of another model. By using larger models with carefully designed prompts that specify how tools should be used, you can generate a high-quality dataset that demonstrates correct tool usage. You can then fine-tune a smaller, faster, and cheaper model to reproduce similar quality.

|                                                                                                  | Training Set Generation                                                                                                | Fine-Tuned Model                                                                        |
| ------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| Model                                                                                            | Large Model                                                                                                            | Smaller Model                                                                           |
| Prompt                                                                                           | Long prompt detailing tool calling strategy. This includes when each tool should be used and which parameters to pass. | Short prompt focused on task. Does not need to address tool-calling strategy in detail. |
| Reasoning Mode                                                                                   | Recommended to Enable                                                                                                  | Optional                                                                                |
| Cost per Token                                                                                   | Expensive                                                                                                              | Cheap                                                                                   |
| Speed                                                                                            | Slower                                                                                                                 | Faster                                                                                  |
| <p>Tool Usage Evals<br><a href="#evaluating-tool-use"><em>Always measure to confirm</em></a></p> | High quality                                                                                                           | High quality                                                                            |
| Origin of Tool Calling Logic                                                                     | Base model + detailed strategy in prompt                                                                               | Learned during fine-tuning                                                              |

#### Create a Tool Calling Fine-Tune

Once you’ve created a training set, return to the “Create a Fine Tune” screen to start training. Select a base model, the tools you want, and the dataset you’ve created to start training. See our [fine-tuning guide](/docs/fine-tuning/fine-tuning-guide) for additional details.

<figure><img src="/files/Ql6fGF4CoXSbrAGezLp7" alt="" width="375"><figcaption></figcaption></figure>

{% hint style="info" %}
You must select a base model that supports tool calling. Kiln will disable tool calling training if the selected base model was not trained for tool calling.
{% endhint %}

{% hint style="success" %}
Kiln will convert your training data into the base model's tool calling format automatically.
{% endhint %}

#### Running a Tool Calling Fine-Tune

When running a Tool Calling Fine-Tune in Kiln, we’ll automatically populate the same set of tools it was trained on.

Adding or removing tools will show a warning, as this model is unlikely to perform well with tools that were not in its training dataset.

{% hint style="info" %}
If deploying a fine-tune created in Kiln, always provide the same tools as it was trained to use.
{% endhint %}

#### Evaluating Tool Use

Kiln has a custom eval for [evaluating appropriate tool use](/docs/evals-and-specs/evaluate-appropriate-tool-use). We suggest using it to compare the base model to your fine-tune, to confirm performance is improving.


# Evals

Powerful tools to measure and optimize your AI systems

#### Two Ways to Build Evals

Kiln has two powerful ways to build evals that ensure your AI systems perform as expected, drive optimizations and don't regress in quality:

* [**Manual Evals**](/docs/evals-and-specs/evaluations): Build industry standard evals with methods like LLM-as-Judge and G-Eval, or fast deterministic checks like pattern matching and tool call checks.
* [**Eval Builder**](/docs/evals-and-specs/specifications)**:** A guided interactive flow that includes synthetic evaluation data generation, edge case detection, judge prompt generation, and more. It's an easy, fast and more comprehensive way to build evals.

<table data-search="false"><thead><tr><th valign="middle"></th><th valign="middle">Manual Evals</th><th valign="middle">Eval Builder</th><th data-hidden></th></tr></thead><tbody><tr><td valign="middle"><p><strong>LLM-as-Judge</strong></p><p><em>including G-Eval</em></p></td><td valign="middle">✅</td><td valign="middle">✅</td><td></td></tr><tr><td valign="middle"><p><strong>Programmatic Judges</strong></p><p><em>exact match, patterns, tool calls, code</em></p></td><td valign="middle">✅</td><td valign="middle">Coming soon</td><td></td></tr><tr><td valign="middle"><strong>Judge Prompt Creation</strong></td><td valign="middle">Manual</td><td valign="middle">Automatic</td><td></td></tr><tr><td valign="middle"><strong>Edge Case Discovery</strong></td><td valign="middle">Manual</td><td valign="middle">Automatic</td><td></td></tr><tr><td valign="middle"><strong>Eval Data Creation</strong></td><td valign="middle"><p>Manual</p><p><em>With synthetic tooling</em></p></td><td valign="middle">Automatic</td><td></td></tr><tr><td valign="middle"><strong>Eval Accuracy</strong></td><td valign="middle">Variable</td><td valign="middle"><p>High</p><p><em>Human in the loop validation and refinement</em></p></td><td></td></tr><tr><td valign="middle"><strong>Approx. Effort</strong></td><td valign="middle">30 mins+</td><td valign="middle">5-10 mins</td><td></td></tr><tr><td valign="middle"><strong>Needed Expertise</strong></td><td valign="middle">Data Science Basics<br><em>Understand Golden sets, data labeling</em></td><td valign="middle">No experience necessary<br><em>Fully Guided UI</em></td><td></td></tr><tr><td valign="middle"><strong>Kiln Account</strong></td><td valign="middle">Optional</td><td valign="middle">Required</td><td></td></tr><tr><td valign="middle"><strong>Docs</strong></td><td valign="middle"><a href="/pages/2PbHtsJUdtp2xiIqeHqQ">Evals Guide</a></td><td valign="middle"><a href="/pages/DM6BHkPtRZJ6F97E5ggL">Eval Builder Guide</a></td><td></td></tr></tbody></table>

#### Guides

* [Eval Builder Guide](/docs/evals-and-specs/specifications): build an eval, synthetic data, and align your judge in one interactive flow
* [Evals 101](/docs/evals-and-specs/evaluations): build your first eval start to finish
* [Judge Types](/docs/evals-and-specs/judge-types): all of Kiln's judge types, from LLM as Judge to deterministic checks, and how to pick between them
* [LLM Judges](/docs/evals-and-specs/llm-judges): grade output with a model and a rubric, including G-Eval and judge prompt guidance
* [Code Judges](/docs/evals-and-specs/code-judges): score an eval with a custom Python function, which can call tools and LLMs of its own
* [Many Small Evals Beat One Big Eval](https://kiln.tech/blog/you_need_many_small_evals_for_ai_products): Blog post which walks through how to setup eval tooling, and how to create an eval culture on your team.
* [Evaluate RAG Accuracy](/docs/evals-and-specs/evaluate-rag-accuracy-q-and-a-evals): Kiln can generate custom Q\&A evals which test your RAG with knowledge from your documents
* [Evaluate Tool Use](/docs/evals-and-specs/evaluate-appropriate-tool-use): ensure your agents are using the right tools, at the right time, with the right parameters with tool use evals
* [Use Kiln Evals on External Agents](/docs/tools-and-mcp/connect-to-existing-agents): If you've built agents in another platform, you can still evaluate them in Kiln using our MCP connectors.


# Evaluations

Evaluate the quality of your models/tasks using state-of-the-art evals

<figure><img src="/files/CxEtDysGyLhtfaO9XOea" alt=""><figcaption></figcaption></figure>

{% hint style="success" %}
The [**Kiln Eval Builder**](/docs/evals-and-specs/specifications) is an even easier way to build evals. It reduces 6 manual steps into one interactive flow.

<a href="/pages/DM6BHkPtRZJ6F97E5ggL" class="button secondary">Try the Kiln Eval Builder</a>
{% endhint %}

### Overview

Kiln includes a complete platform for ensuring your tasks/models are of the highest possible quality. It includes:

* Access a range of judge types: LLM as Judge and G-Eval for subjective quality, plus fast deterministic checks for anything you can state as a rule
* Compare and benchmark your judges against human evals to find the best possible evaluator for your use case
* Test a variety of different methods of running your task (prompts, models, fine-tunes) to find which perform best
* Easily manage datasets for eval sets, golden sets, human ratings through our intuitive UI, including automatic synthetic data generation.
* Generate evaluators automatically. Using your task definition we'll create an evaluator for your task's overall score and task requirements
* Utilize built-in eval templates for toxicity, bias, jailbreaking, and other common eval scenarios
* Integrate evals with the rest of Kiln: use synthetic data generation to build eval sets, or use evals to evaluate fine-tunes
* Optional: Python Library Usage

{% hint style="success" %}
New to evals? We suggest reading our blog post [Many Small Evals Beat One Big Eval](https://kiln.tech/blog/you_need_many_small_evals_for_ai_products), which walks you through how to setup eval tooling, and how to create an eval culture on your team.
{% endhint %}

### Video Guide

{% hint style="info" %}
The UI has been updated since this video was recorded, but the general flow remains the same. Some steps like tagging eval data are now automated for you.
{% endhint %}

{% embed url="<https://vimeo.com/1067948856?share=copy>" %}
Video walkthrough of creating a LLM evaluator
{% endembed %}

### Concepts Overview

This is a quick summary of all of the concepts in creating evals with Kiln:

* Eval (aka Evaluator): defines an evaluation goal (like "overall score" or "toxicity"), and includes dataset definitions to use for running this eval. You can add many evals to a task, each for different goals.
* Score: an output score for an eval like "overall score", "toxicity" or "helpfulness". An eval can have 1 or more output scores. These have a score type: 1-5 star, pass/fail, or pass/fail/critical.
* Judges: methods of running an Eval. A judge includes a [judge type](/docs/evals-and-specs/judge-types), and whatever that type needs to run: judge instructions and a model/provider for an LLM judge, or a rule like a regular expression or an expected set of tool calls for a programmatic judge. An eval can have many judges of different types, and Kiln will help you compare them to find which best correlates to human preferences.
* Task Run Methods: methods of running your task. A task run method includes a prompt, model, model provider and options (temperature, top\_p, etc). A task can have many run methods. Once you have an Eval, you can use it to find an optimal run-method for your task: the run method which scores the highest, using your eval.

### The Workflow

Working with Evals in Kiln is easy. We'll walk through the flow of creating your first evaluator end to end:

* [Creating an Evaluator](#creating-an-eval)
* [Add a Judge to your Eval](#add-a-judge-to-your-eval)
* [Create your Eval Datasets](#create-your-eval-datasets)
* [Finding the Ideal Judge](#finding-the-ideal-judge)
* [Finding the Ideal Run Method](#finding-the-ideal-run-method)
* [Iterate and Expand](#iterate-and-expand)

<figure><img src="/files/sB00FZfWJJwBEt9IX0tW" alt="" width="188"><figcaption><p>Kiln's UI will guide you</p></figcaption></figure>

### Creating an Eval

From the "Evals" tab in Kiln's UI, you can easily create a new evaluator.

#### Pick a Goal / Select a Template

Kiln has a number of built-in templates to make it easy to get started.

{% hint style="info" %}
We recommend starting with the "Overall Score and Task Requirements" template and "Issue" template for bugs.
{% endhint %}

* **Overall Score and Task Requirement Scores:** Generate scores for the requirements you set up when you created this task, plus an overall-score. These can be compared to human ratings from the dataset UI.
* **Kiln Issue Template**: evaluate an issue or bug you've seen in your task. You'll describe the issue and provide examples. Kiln will help generate synthetic data to reproduce the issue which can help you ensure your fix works. For advanced issues, Kiln can generate synthetic training data for fine-tuning a model to avoid this issue.
* **Built-in Templates**: Kiln includes a number of common templates for evaluating AI systems. These include evaluator templates for measuring **toxicity, bias, maliciousness, factual correctness, and jailbreak susceptibility**.
* **Custom Goal and Scores**: If the templates aren't a good fit, feel free to create your own Eval from scratch using the custom option. However, prefer the "issue" template where possible as it's integrated into synthetic data generation.

Select a template, edit if desired, and save your eval.

### Add a Judge to your Eval

The Eval you created defines the goal of the eval, but it doesn't include the specifics of how it's run. That's where judges come in — they define the exact approach of running an eval. This includes things like the judge type, and for LLM judges the judge model/provider and judge prompt.

#### Select a Judge Type

The judge type defines how your eval is actually scored. Kiln offers two families:

* **LLM Judges**: a model reads the output and grades it against a rubric you write. Best for subjective qualities like tone, helpfulness, toxicity, etc. See our [LLM Judges](/docs/evals-and-specs/llm-judges) guide for the options you can configure.
* **Programmatic Judges**: code inspects the output or the trace and returns a pass/fail — exact matches, regular expressions, tool call trajectories, step counts, or a custom Python function. There's no model call, so they're fast, free, and return the same answer every time.

If your eval goal can be stated as a rule, prefer a programmatic judge. If it needs judgement, use an LLM judge. See our [Judge Types](/docs/evals-and-specs/judge-types) guide for all of the options, and help choosing between them.

### Create your Eval Datasets

An eval in Kiln has several datasets, each defining a subset of the items in your task's dataset:

* **Test Dataset**: held-out data for measuring final quality. This is the data used when evaluating different methods of running your task, and the scores shown in the "Compare" view. Every eval needs one.
* **Golden Dataset**: the data used when trying to find the best judge for this eval. These items have human ratings, so we can compare judges to human preference.
* **Training Dataset** \[optional]: used by optimizers, such as the [automatic prompt optimizer](/docs/prompts/automatic-prompt-optimizer).
* **Validation Dataset** \[optional]: also used by optimizers, to confirm a result generalizes beyond the data it was tuned on.

This section will walk you through populating your eval datasets.

#### Defining your Dataset with Tags

When first creating your eval, you will specify a "tag" which defines each of these datasets as a subset of all the items in Kiln's Dataset tab. To add/remove items from your datasets, simply add/remove the corresponding tag. These tags can be added or removed anytime from the "Dataset" tab.

Don't worry if your dataset is empty when creating your eval, we'll guide you through adding data after its creation.

By default, Kiln will suggest appropriate tags and we suggest keeping the defaults. Tags are named after your eval: an eval named "toxicity" gets the tags "test\_toxicity", "golden\_toxicity", "train\_toxicity" and "val\_toxicity". Evals created by earlier versions of Kiln keep the tags they were created with.

{% hint style="info" %}
"Golden" is a term often used in data science, to describe a "gold standard" dataset, used to compare different methods/approaches.
{% endhint %}

{% hint style="info" %}
If you're creating multiple evals for task, it's usually beneficial to maintain separate datasets for each eval. For example, a "toxicity" eval dataset will likely be filled with negative content you wouldn't want in your overall-score eval. Kiln will suggest goal-specific tags by default.
{% endhint %}

#### Populating the Dataset with Synthetic Data

Most commonly, you'll want to populate the datasets using synthetic data. Clicking `Add Eval Data` then `Synthetic Data` from the Evals UI; this will launch the synthetic data gen tool with the proper [eval tags](#defining-your-dataset-with-tags) already populated. See our [synthetic data generation guide](/docs/synthetic-data-generation) for details on generating synthetic data.

We suggest at least 160 data samples per eval. Difficult or subjective tasks may require more.

An appropriate data gen template will be populated when you enter data-gen via an eval. You can customize this template to guide data generation. See [the docs](/docs/synthetic-data-generation#templates-and-custom-guidance) for details.

{% hint style="info" %}
Golden eval datasets work best if they have a range of ratings (some pass, some fail, some of each star-score).

If your dataset doesn't have enough variation, you may see "N/A" scores when comparing evaluators.

If after rating your golden set doesn't a range of content (for example, one score always passes or always fails), generate some additional content for the missing cases. You can use human guidance to do this, see "Guidance Templates" below for examples.
{% endhint %}

#### Tagging the Eval Datasets

If you've launched the synthetic data generator from the evals UI, it will know to assign the needed tags in the appropriate ratio. There's no need to manually tag dataset items.

<figure><img src="/files/ZXjCxPopPEG9Myx3fbKn" alt=""><figcaption><p>Data gen will populate target tags when launched from Evals</p></figcaption></figure>

<details>

<summary>Adding tags manually</summary>

If your dataset items weren't automatically tagged for any reason, you can also add tags manually:

1. Add your data to Kiln using one of the import options ([CSV import](/docs/organizing-datasets#importing-data-into-your-dataset), [python library import](/developers/python-library-quickstart))
2. Open the "Dataset" tab in kiln
3. Filter your dataset to only the content you want to tag. For example, synthetic data is tagged with an automatic tag such as synthetic\_session\_12345, CSV imports have similar tags.
4. Use the "Select" UI to select a portion of your dataset for your eval-dataset. 80% is a good starting point. Add the tag for your test dataset, which is "test\_\<eval name>" if you kept the default tag names. Note: if you generated data using synthetic "topics", make sure to include a mix of each topic in each sub-dataset.
5. Select only the remaining items, and add the tag for your golden dataset, which is "golden\_\<eval name>" if you kept the default tag names.
6. Filter the dataset to both tags (your test tag and your golden tag) to double check you didn't accidentally add any items to both datasets.

</details>

{% hint style="info" %}
**Training and Validation Sets**

If you plan to use an optimizer like the [automatic prompt optimizer](/docs/prompts/automatic-prompt-optimizer), add training and validation datasets too. Keeping them separate from your test dataset is what stops an optimizer from tuning against the same data you use to measure the result.

You can create these now, or generate them later — they're optional until an optimizer needs them.
{% endhint %}

#### Add Human Ratings

Next we'll add human ratings, so we can measure how well our eval judge performs compared to a human. If you have a subject matter expert for your task, get them to perform this step. See our [collaboration guide](/docs/collaboration) for how to work together on a Kiln project.

The `Rate Golden Dataset` button in the eval screen will take you to the dataset view filtered to your golden dataset (the items which need ratings). Once fully rated, this will get a checkmark and you can proceed to the next step

<figure><img src="/files/Rhtw33iLEB5gtuCQ9JPd" alt="" width="375"><figcaption></figcaption></figure>

{% hint style="success" %}
You can use the left/right keyboard keys to quickly move between items. Only the golden dataset needs ratings, not the test dataset.
{% endhint %}

### Finding the Ideal Judge

{% hint style="info" %}
**Who Judges the Judge?**

While it is relatively easy to create a LLM-as-Judge eval, an important question remains — does it actually work?

In this section we use a human judge's ratings to ensure our LLM-as-Judge aligns to human ratings, so we have trust in our system.
{% endhint %}

You added a Judge to your eval above. However, we don't actually know how well this judge works. Kiln includes tools to compare multiple judges, and find which one is the closest to a real human evaluator.

It may seem strange, but yes… one of the first steps of building an eval is to judge judges (not a typo). It sounds complicated, but Kiln makes it easy.

#### Run Evals on your Golden Set

Open your eval from the "Evals" tab, then click the "Compare Judges" button. From the "Compare Judges" screen, click the "Run Eval" button.

This will run your eval on the golden dataset, once with each judge.

Once complete, you'll have a set of metrics about how well the judge's scoring matched human scores.

#### Add Judges and Compare to Find the Best

One score in isolation isn't helpful. You'll want to add additional judges to see which one performs best. Kiln makes it easy to compare judges. We suggest trying a range of options:

* Try a range of different models: you may be surprised which model works best as an evaluator for your task. Be sure to try SOTA models, like the latest models from OpenAI and Anthropic. Even if you prefer open models, it can be good to know how far you are from these benchmarks.
* Try custom eval instructions, not just the template contents.

Once you've added multiple judges, you can compare scores to find the best evaluator for your task. You're looking for the score which appears highest in the table, which means the least deviation from human scores. On some scoring methods higher scores are better (Kendall's, Spearman) and on others lower is better (MSE, MAE); the table will be sorted so the best are at the top.

#### Understanding Correlation Scores

There's no benchmark good/bad score for an evaluator; it all depends on your task difficulty.

For an easy and highly deterministic task, you might be able to find many judges which achieve near perfect scores, even with small eval models and default prompts.

For a highly subjective task, it's likely no evaluator will perfectly match the human scores, even with SOTA models and custom prompts. It's often the case that two humans can't match each other on subjective tasks. Try a range of judges, and pick the one with the best score.

The more subjective the task, the more beneficial a larger and more diverse golden dataset becomes.

<details>

<summary>Technical comparison of score options: Kendall' Tau, Spearman, Pearson, Mean Squared Error, Mean Absolute Error</summary>

> Each score is a correlation score between the judge's scores and the human scores.

**TL;DR**

We suggest you use Kendall Tau correlation scores to compare results.

Kendall Tau scores range from -1.0 to 1, with higher values being higher correlation between the human ratings and the automated judge's scores.

The absolute value of Kendall Tau scores will vary depending on how subjective your task is. Find the highest score for your task, and select it as your default judge.

**Spearman, Kendall Tau, and Pearson Correlation**

*From -1 to 1. Higher is better.*

These are three scientific correlation coefficients. For all three, the value tends to be high (close to 1) for samples with a strongly positive correlation, low (close to -1) for samples with a strongly negative correlation, and close to zero for samples with weak correlation. Scores may be 'N/A' if there are too few samples or not enough scoring variation in your human-rated dataset (golden data).

* Spearman evaluates the rank of the scores, not the absolute values.
* Kendall's Tau evaluates rank order of pairs. It is more robust to outliers, handles ties better, and performs better on small datasets. As our datasets often have ties (pass/fail and 5-star datasets have limited discrete values), we suggest Kendall's Tau.
* Pearson evaluates linear correlation.

**Mean Absolute Error**

*Lower is better*

Example: If a human scores an item a 3, and the eval scores it a 5, the absolute error would be 2 \[abs(3-5)]. The overall score is the mean of all absolute errors.

**Normalized Mean Absolute Error**

*Lower is better*

Like mean absolute error, but scores are normalized to the range 0-1. For example, for a 1-5 star rating, 1-star is score 0 and 5-star is score 1.

**Mean Squared Error**

*Lower is better*

Example: If a human scores an item a 3, and the eval scores it a 5, the squared error would be 4 \[(3-5)^2]. The overall score is the mean of all squared errors. This improves over absolute error as it penalizes larger errors more.

**Normalized Mean Squared Error**

*Lower is better*

Like mean squared error, but scores are normalized to the range 0-1. For example, for a 1-5 star rating, 1-star is score 0 and 5-star is score 1.

</details>

<details>

<summary>Resolving "N/A" Correlation Scores</summary>

If you see "N/A" scores in your correlation table, it means more data is needed. This can be one of two cases

* ***Simply not enough data***: if your golden dataset is very small (<10 items) it can be impossible to produce confident correlation scores. Add more data to resolve this case.
* ***Not enough variation of human ratings in the golden dataset***: if you have a larger dataset, but still get N/A, it's likely there isn't enough variation in your dataset for the given score. For example, if all of the golden samples of a score pass, the evaluator won't produce a confident correlation score, as it has no failing examples and everything is a tie. Add more content to your golden dataset, designing the content to fill out the missing score ranges. You can use synthetic data gen [human guidance](/docs/synthetic-data-generation#templates-and-custom-guidance) to generate examples that fail.

</details>

#### Select the Winning Judge

Once you have a winner, click the "Set as default" button to make this judge the default for your eval.

<figure><img src="/files/ACHiDmXhaqoXAH7bF1fA" alt="" width="179"><figcaption><p>Select the default judge</p></figcaption></figure>

### Finding the Ideal Run Method

Now that we have an evaluator we trust, we can use it to rapidly evaluate a variety of methods of running our task. We call this a "Run Method" and it includes the model (including fine-tunes), the model provider, and the prompt.

Return to the "Evaluator" screen for your eval, and add a variety of run methods you want to compare. We suggest:

* A range of models (SOTA, smaller, open, etc)
* A range of prompts: both Kiln's [auto-generated prompts](/docs/prompts#prompt-generators), and [custom prompts](/docs/prompts#custom-prompts-saved-prompts)
* A range of model parameters: temperature, top\_p, etc
* Some model fine-tunes of various sizes, created by [Kiln fine tuning](/docs/fine-tuning/fine-tuning-guide)

Once you've defined a set of run methods, click "Run Eval" to kick off the eval. Behind the scenes, this is performing the following steps:

* Fetching the input data from your eval's test dataset
* Generating new output for each input, using each run method you defined for each input
* Running your evaluator on each result, collecting scores

#### Comparing Run Methods

Once done, you'll have results for how each run method performed on the eval.

These results are easy to interpret compared to the judge comparisons. Each score is simply the average score from that run method. Assuming we want to find the run method that produces the best content, simply find the highest average score.

Congrats! You've used systematic evals to find an optimal method for running your task!

### Comparing Run Methods Over Many Evals

In the last step, you found the ideal run method for a specifc eval. However, over time your team will generate many evals.

When you want to try a new model or prompt, you'll want to make sure the new method is better, not just on a single eval, but across all prior evals.

Kiln's compare view makes it easy to compare run methods across many evals. It also lets you compare the cost difference of each method:

<figure><img src="/files/VYDOyZZVu6q0mZJkdlkg" alt="" width="375"><figcaption><p>Comparing several run methods across all evals</p></figcaption></figure>

Click "Compare" in the "Evals" tab to launch this feature:

<figure><img src="/files/ZEUyAlpkbzolPWI33V2n" alt="" width="375"><figcaption></figcaption></figure>

### Philosophy: AI Product Evals work Best with Many Small Evals <a href="#setup-team-evals" id="setup-team-evals"></a>

At Kiln we believe if creating an eval takes less than 10 minutes, your team will create them when they spot issues or fix bugs.

When evals become a habit instead of a chore, your AI system becomes dramatically more robust and your team moves faster.

Read our ~~manifesto~~ [guide on how to setup evals for your team](https://kiln.tech/blog/you_need_many_small_evals_for_ai_products#setup-team-evals). It covers:

* Many Small Evals Beat One Big Eval, Every Time
* The Benefits of Many Small Evals
* Evals vs Unit Testing
* 3 Steps to Set Up Your Team for Evals and Iteration

### Next Steps: Iterate and Expand

Congrats, you've found an optimal method of running your task!

However there's always room to improve.

#### Iterate on Methods

You can repeat the processes above to try new judges or run-methods. Through more searching, you may be able to find a better method and improve overall performance.

You can iterate by trying new prompts, more models, building custom fine-tuned models, or trying new state of the art models as they are released.

#### Expand your Dataset

Your understanding of your model/product usually gets better over time. Consider adding data to your dataset over time (both test and golden). This can come from real users, bug reports, or new synthetic data that comes from a better understanding of the problem. As you add data, re-run both sub-evals (judge and run-method) to find the best judge and run-method for your task.

#### Add New Evals

You can always add additional evals to your Kiln project/task. Try some of our built-in templates like [issue evals](/docs/issues), bias, toxicity, factual correctness, or jailbreak susceptibility — or create your own from scratch!

Most commonly, you'll collect a list of ["Issue" evals](/docs/issues) over time. This set of evals helps you work with confidence that new changes aren't regressing old issues.

Read our blog [Many Small Evals Beat One Big Eval, Every Time](https://kiln.tech/blog/you_need_many_small_evals_for_ai_products) for a quality strategy that scales as your product and team grow.

### Optional: Python Library Usage

For developers, it's also possible to use evals from our [python library](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html).

Be aware, in our library task run methods are called TaskRunConfigs and judges are called EvalConfigs.

See the EvalRunner, Eval, EvalConfig, EvalRun, and TaskRunConfig classes for details.


# Eval Builder

The Kiln Eval Builder combines evals, synthetic data generation, automatic judge prompt creation, and edge case detection into one easy to use feature

### Demo & Quick Start

{% embed url="<https://vimeo.com/1161246105>" %}

{% hint style="info" %}
**Note:** The Eval Builder requires a Kiln Pro account. Registration is free and easy inside the Kiln app.
{% endhint %}

{% hint style="info" %}
The Eval Builder currently only supports the creation of LLM as Judge evals.
{% endhint %}

### What is the Kiln Eval Builder?

<figure><img src="/files/8IWCv4IiKmYo1qwIDTY9" alt="" width="375"><figcaption></figcaption></figure>

The Kiln Eval Builder combines Kiln’s best features into one interactive tool: evals, synthetic data generation, automatic judge prompt creation, and edge case detection. Together they go beyond making an eval manually in several ways:

* **Identify Gaps with AI**: Kiln will read your judge prompt and help refine it. We detect underspecified aspects of your judge, conflicts with your task definition, ambiguous aspects that Judges may struggle with, and other common issues. It then works with you to close gaps and refine conflicts.
* **Interactive Human Alignment & Accuracy**: Building a LLM-as-Judge as good as a human isn’t easy. Human judges make subtle and subjective decisions, and have a hard time articulating their judgement process in a way LLMs can duplicate. Our alignment loop finds tough edge cases, compares LLM judge to human preference, and works with you iteratively until your judge is aligned to your preference.
* **Automatic Synthetic Data**: Build robust synthetic dataset generator as you work. By the time you save your eval you’ll have large and accurate datasets for evals and training.
* **Judge Meta-prompting**: Humans often struggle at writing effective eval judge prompts. Our judge meta-prompting takes your issues and concerns in human terms, and turns them into accurate and judge-able evals.
* **Easy to Use**: Subject matter experts can easily create accurate evals, without a lengthy iteration loop with data scientists. The Eval Builder will walk you through all the steps of defining your judge, creating synthetic data, golden dataset, aligning your judge, and creating training datasets. You get the same rigorous process, without managing each step.
* **Fast**: creating an eval can be done in as little as 5 minutes, compared to over 30 minutes manually.

### How to Get Started

Getting started is easy:

* Open the Kiln App to any task
* Click "Evals" in the sidebar
* Click "Create Eval"
* Connect Kiln Pro account (if you haven’t already)
* Follow the interactive steps until complete!

See the video above for a complete walkthrough.


# Judge Types

Score your evals with LLM judges, or with fast deterministic checks that cost nothing to run

A judge is a method of running an eval: it takes the output of a task run, and produces scores. Every eval in Kiln can have many judges, and each judge has a type.

Kiln offers two families of judge types:

* **LLM Judges**: a model reads the output and grades it against criteria you write. Best for subjective qualities like tone, helpfulness, toxicity, etc.
* **Programmatic Judges**: code inspects the output or the trace and returns a pass/fail. No model call, so they're fast, free, and return the same answer every time. Best for anything you can state as a rule.

{% hint style="success" %}
**Prefer a programmatic judge when one fits.**

An LLM judge that decides "did the agent call `get_weather` before answering?" costs money on every run and can change its mind. A [Tool Call Check](#tool-call-check) answers the same question for free, identically, every time.

Save LLM judges for the questions that genuinely require subjective judgement.
{% endhint %}

### The Judge Types

| Judge Type                            | Family             | What it scores                                                                 |
| ------------------------------------- | ------------------ | ------------------------------------------------------------------------------ |
| [LLM as Judge](#llm-as-judge)         | LLM Judge          | A model grades the output against a rubric you write                           |
| [Code](#code-beta)                    | Programmatic Judge | A custom Python `score()` function you write                                   |
| [Tool Call Check](#tool-call-check)   | Programmatic Judge | The agent called the right tools, in the right order, with the right arguments |
| [Exact Match](#exact-match)           | Programmatic Judge | The output equals an expected value                                            |
| [Pattern Match](#pattern-match)       | Programmatic Judge | The output matches (or doesn't match) a regular expression                     |
| [Contains](#contains)                 | Programmatic Judge | The output contains (or omits) a substring                                     |
| [Set Check](#set-check)               | Programmatic Judge | A set of values from the output matches an expected set                        |
| [Step Count Check](#step-count-check) | Programmatic Judge | The agent finished within an expected number of steps                          |

To add a judge, open your eval and create a new judge, or pick a type directly from the "Select an Eval Type" screen when creating a new eval. Programmatic judges are listed under the "Programmatic Judges" heading.

### Choosing a Judge Type

Work down this list and stop at the first one that fits:

1. **Is there exactly one right answer?** Use [Exact Match](#exact-match) (or [Set Check](#set-check) if the answer is a set of values).
2. **Can you state the rule as text matching?** Use [Contains](#contains) or [Pattern Match](#pattern-match). Good for format rules ("always ends with a citation", "never mentions a competitor").
3. **Is it about what the agent did, not what it said?** Use [Tool Call Check](#tool-call-check) for which tools ran, or [Step Count Check](#step-count-check) for how many steps it took.
4. **Is the rule real, but too complex for the above?** Use a [Code](#code-beta) judge. It can also call an LLM for the subjective part, so you can filter a huge trace down in code and only pay for a judgement on what's left. See [Code Judges](/docs/evals-and-specs/code-judges).
5. **Otherwise, use an** [**LLM as Judge**](#llm-as-judge). Subjective quality needs a model.

{% hint style="info" %}
A single eval can have judges of different types, and Kiln will show their scores side by side. A "helpfulness" eval can pair an LLM judge for tone with a Pattern Match that catches a formatting bug — you don't need to pick just one.
{% endhint %}

### Output to Check

The Exact Match, Pattern Match, Contains and Set Check types ask which part of the run they should look at:

* **Final Message**: the model's final output. The default, and what you want most of the time.
* **Entire Trace**: the whole conversation, including tool calls, as JSON.
* **Custom (Jinja)**: extract part of the output or trace using a Jinja expression.

Custom expressions are useful for structured outputs and agent traces. Some examples:

| Goal                               | Expression                                |
| ---------------------------------- | ----------------------------------------- |
| Extract a field from JSON output   | `(final_message \| fromjson).user.status` |
| Truncate a long output             | `final_message \| truncate(200)`          |
| Last message in the trace          | `trace[-1].content`                       |
| Count messages in the trace        | `trace \| length`                         |
| Name of a tool called in the trace | `trace[-1].tool_calls[0].function.name`   |

Tool Call Check and Step Count Check always read the trace, so they don't offer this option. [Code](#code-beta) judges receive the output and trace directly, and pick out what they need in Python.

### LLM as Judge

A model reads the output and grades it against criteria you write, producing a score for each of your eval's output scores. You choose the judge model and provider, write the judge prompt and evaluation instructions, and can optionally enable G-Eval for more nuanced scores.

See our [LLM Judges](/docs/evals-and-specs/llm-judges) guide for each of these options.

### Programmatic Judges

#### Exact Match

Passes when the output equals an expected value.

* **Expected Value**: the value the output should equal.
* **Case Sensitive**: on by default.
* **Output to Check**: see [above](#output-to-check).

Best for tasks with a single correct answer: classification labels, yes/no answers, extracted IDs.

#### Pattern Match

Passes when the output matches a regular expression — or when it doesn't, if you set the mode to "must not match".

* **Expected Pattern (Regex)**: any Python regular expression.
* **Match Mode**: must match, or must not match.
* **Output to Check**: see [above](#output-to-check).

Best for format rules. A "must not match" pattern is a cheap way to catch an output that keeps leaking something it shouldn't, like a placeholder string or an internal ID format.

#### Contains

Passes when the output contains a substring — or when it doesn't, if you set the mode to "must not contain".

* **Expected Substring**: the text to look for.
* **Case Sensitive**: on by default.
* **Match Mode**: must contain, or must not contain.
* **Output to Check**: see [above](#output-to-check).

Simpler than Pattern Match, and usually clearer to a teammate reading your eval later. Reach for Pattern Match only when a plain substring won't do.

#### Set Check

Parses a set of values from the output and compares it to an expected set.

* **Expected Values**: the values to compare against.
* **Comparison Mode**: `subset` (everything found must be in the expected set), `superset` (everything expected must be found), or `equal` (exactly the same values).
* **Output to Check**: see [above](#output-to-check).

Best for multi-label classification, tag extraction, or any task where the output is a list and the order doesn't matter.

#### Tool Call Check

Inspects the agent's trace to check it called the tools you expected.

* **Expected Tools**: one or more tools, each optionally with expected arguments. Each argument can be matched exactly, by substring, or by regular expression.
* **Match Mode**:
  * All (any order): every expected tool was called
  * Any: at least one expected tool was called
  * Ordered (in list order): the expected tools were called, in the order listed
  * Never: none of the listed tools were called
* **Unlisted Tool Calls**: allow any other tool the agent calls, or fail the check if it calls anything not on your list. Not shown when Match Mode is "Never".

This is the deterministic way to test tool use. It answers "did the agent do the right thing?" for free, where an LLM judge would need to read the whole trace and form an opinion. See [Evaluate Appropriate Tool Use](/docs/evals-and-specs/evaluate-appropriate-tool-use) for the full workflow, including when you still want an LLM judge.

#### Step Count Check

Counts steps in the agent's trace and passes when the count is within bounds you set.

* **What to Count**: tool calls, model responses, or conversation turns.
* **Bounds**: a **Minimum**, a **Maximum**, or both — at least one is required.

Best for agent efficiency. Cap an agent at 5 tool calls to catch runaway loops, or require at least 1 to confirm it actually used a tool instead of answering from memory.

#### Code (Beta)

Write a custom Python `score()` function. It can read the output, the full trace, and the task input, and it can call other tools — including built-in tools that call an LLM.

See the [Code Judges](/docs/evals-and-specs/code-judges) guide for details.

### Judges and Human Ratings

Kiln's [judge comparison](/docs/evals-and-specs/evaluations#finding-the-ideal-judge) tools measure how closely a judge's scores match human ratings from your golden dataset. This is essential for LLM judges: an LLM judge is an approximation of human preference, and you need to know how good the approximation is.

Programmatic judges are a different kind of thing. They don't approximate anything — a Pattern Match either encodes the rule you meant or it doesn't, and it will return the same answer forever. You generally don't need a golden set or human ratings to trust one.

Running judge comparison on a programmatic judge is still allowed, and there's one case where it's genuinely useful: confirming that the rule you wrote actually captures what humans care about. If your Tool Call Check passes items your subject matter expert would fail, the check is wrong, not the human — and comparison will show you that.

{% hint style="success" %}
If your eval only uses programmatic judges, you can skip the golden dataset and human rating steps entirely, and go straight to [comparing run methods](/docs/evals-and-specs/evaluations#finding-the-ideal-run-method).
{% endhint %}


# LLM Judges

Grade your task's output with a model and a rubric you write

An LLM judge uses a model to grade the output of your task, against criteria you write. It combines a "thinking" stage (chain of thought/reasoning), followed by asking the model to produce scores matching the goals you laid out in your eval.

LLM judges are the right tool for subjective qualities like tone, helpfulness, toxicity, etc. For anything you can state as a rule, a [programmatic judge](/docs/evals-and-specs/judge-types#programmatic-judges) will be faster, cheaper, and perfectly consistent.

This guide covers the options you'll configure when adding an LLM judge. For the overall eval workflow, see [Evaluations](/docs/evals-and-specs/evaluations).

### Select a Judge Model & Provider

Select the model you want the judge to use (including which AI provider it should be run on).

{% hint style="info" %}
We suggest larger high quality models for judges, as you'll be trusting their results to make product improvements. You can always run a cheaper/smaller model for inference which is where the majority of compute is spent in most projects.
{% endhint %}

<details>

<summary><strong>Using models to evaluate models? Does that really work?</strong></summary>

Your intuition might be that you can't use LLMs to evaluate LLM tasks. Won't they make the same errors during evaluation that they make running your task?

There's a few reasons this approach actually works quite well:

* You can use better/larger models during evaluations: evals are (typically) run less often than the task itself. You can use larger models during evals, to gain trust in your smaller/faster task model.
* You can use more inference time compute during evaluation. Evals can be run with advanced reasoning models or detailed chain-of-thought instructions during eval, since latency and cost matter less during evals (they are run less often, and your users aren't waiting for an answer). In particular, defining specialized eval prompts covering specific error cases to watch for, multi-shot examples and rating guidance can really help evals outperform the core task model.
* Often we see the best model at evaluating a task is not the best model at running the task. Using the best model for each job can improve your overall system.

</details>

### G-Eval

G-Eval is an enhanced form of LLM as Judge. It looks at token output probabilities (logprobs) to create a weighted score. For example, if the model had a 51% chance of passing an eval and 49% chance of failing it, G-Eval will give the more nuanced score of 0.51, where LLM-as-Judge would simply pass it (1.0). The [G-Eval paper (Liu et al)](https://arxiv.org/abs/2303.16634) compares G-eval to a range of alternatives (BLEU, ROUGE, embedding distance scores), and shows it can outperform them across a range of eval tasks.

{% hint style="info" %}
Since G-Eval requires logprobs (individual token probabilities), only a limited set of models and providers work with G-Eval. Currently it only works best with older OpenAI models like GPT 4.1. G-eval is not recommended for reasoning models, as the reasoning colors the probability distribution.

The UI will only show the G-Eval option if you select a supported model + provider.

Unfortunately [Ollama doesn't support logprobs yet](https://github.com/ollama/ollama/issues/2415).
{% endhint %}

### Evaluation Instructions

LLM judges give the model time to "think" using chain-of-thought/reasoning before generating the output scores. Evaluation instructions are an ordered list of steps, giving the model a process for "thinking through" the eval prior to answering.

This section appears when your judge prompt doesn't already carry its own steps — typically an eval you created without a template. If you created your eval from a template, the steps come from that template, and you edit them in the judge prompt itself (see below).

<details>

<summary>Advanced tactics for defining evaluation instructions</summary>

If you start editing the eval's instructions, here are some advanced tactics/guidance that can help improve your eval performance:

* Include Multi-shot examples: for a step, give examples of different outputs and how they should be scored. Be sure to not include examples in your eval datasets.
* If your eval has multiple output scores, include at least 1 step for each score.
* Consider order of your steps: start with the more independent considerations, before moving to holistic considerations. For example, instructions for generating a final "overall score" should come after all other thinking steps.
* Consider short-circuit exits and limits: for example "If this step results in a failure, always return a 1-star overall score." or "If this step fails, the maximum overall score you should return is 3-stars".
* Consider weighting guidance for overall scores: if you have many steps producing an overall score, tell the LLM which steps matter the most.

</details>

### Advanced: Judge Prompt

The judge prompt is the Jinja2 template Kiln uses to prompt the judge model, and it's what actually carries your rubric. Kiln assembles a default for you from your task's description, the data to evaluate, and your evaluation instructions — so editing it is optional.

Open "Advanced: Judge Prompt" to edit it, or to set a **System Prompt** for the judge model. This is also where you edit template-derived evaluation steps, since those arrive as part of the prompt rather than as a separate list.

{% hint style="info" %}
The evaluator model can almost always perform better if it knows what the task was. Kiln includes your task's description in the default judge prompt for exactly this reason — keep that description short and clear, usually a sentence, and your judges benefit along with your task.
{% endhint %}

### Align your Judge to Human Preference

An LLM judge is an approximation of human preference, so you'll want to know how good the approximation is. Kiln compares your judge's scores against human ratings from your golden dataset, and helps you try several judges to find the one that matches your raters most closely.

See [Finding the Ideal Judge](/docs/evals-and-specs/evaluations#finding-the-ideal-judge) for that workflow.


# Code Judges

Score your evals with a custom Python function

{% hint style="info" %}
Code judges are in **beta**. The authoring contract may change in future releases.
{% endhint %}

A code judge scores your eval with a Python function you write, instead of a model or a built-in check. It's the escape hatch for rules that are real and checkable, but too complex for the other [judge types](/docs/evals-and-specs/judge-types): scoring against a lookup table, parsing a structured output and validating it, or measuring something specific to your domain.

Code judges are fast and cheap compared to [LLM as Judge](/docs/evals-and-specs/judge-types#llm-as-judge), and unlike an LLM judge they return the same score every time.

### Creating a Code Judge

Add a judge to your eval and select the "Code" type.

{% hint style="warning" %}
**Code judges run on your machine, with full access.**

There are no import restrictions or resource limits beyond the timeout. Adding or editing code in a project requires confirming you trust it — that trust gate is the security boundary. See [Code Trust](/docs/tools-and-mcp/code-tools#code-trust) for details.
{% endhint %}

### The `score()` Function

Your code must define a function named `score()`. Kiln passes it only the arguments your function actually declares, so declare the ones you need and omit the rest. You must accept at least `output` or `trace`.

| Parameter        | What you get                                         |
| ---------------- | ---------------------------------------------------- |
| `output`         | The model's final output, as a string                |
| `trace`          | The full conversation, as a list of message dicts    |
| `task_input`     | The original task input, as a string                 |
| `reference_data` | A dict of ground-truth data for this item, or `None` |

`score()` returns a dict, keyed by the JSON key of each of your eval's output scores. A score's JSON key is its display name in lowercase snake\_case: a score named "Exact Match" has the key `exact_match`. Kiln checks the returned keys against exactly this set, so return the JSON keys, not the display names.

```python
def score(output: str) -> dict:
    """Pass when the output is valid JSON with a non-empty summary."""
    import json

    try:
        parsed = json.loads(output)
    except json.JSONDecodeError:
        return {"valid_output": 0.0}

    has_summary = bool(parsed.get("summary", "").strip())
    return {"valid_output": 1.0 if has_summary else 0.0}
```

Scores are floats. For a pass/fail score, return `1.0` or `0.0`.

{% hint style="info" %}
Test your judge before saving it. A test run checks the dict you return against your eval's output scores, so a missing, misspelled or extra key is caught while you're still in the editor — not partway through an eval run.
{% endhint %}

### Calling LLMs from a Code Judge

Code judges can call two built-in tools, selectable from the judge's tool picker like any other tool:

* **`llm`**: a general-purpose model call. Pass a prompt, model and provider. Provide an optional JSON schema to force structured output; without one you get text back.
* **`llm_judge`**: the same, but it automatically applies your eval's own output score schema, so the scores come back keyed the way your eval expects.

Both take a `prompt` that is rendered as a **Jinja2 template** against an `input` dict. Put your instructions in the template and pass the run's data through `input` — never build the prompt by interpolating trace text directly, or a `{{ }}` appearing in that text will be parsed as Jinja and fail the call.

Both run in Kiln itself rather than in your judge's subprocess, so your API keys are never exposed to your code.

This combination solves the long-trace problem. An agent run can produce a 500k token trace, and handing all of it to an LLM judge is slow and expensive. Instead, filter it down in code to the handful of messages that actually matter, then ask a cheap model about just those:

```python
import json
from kiln import tools


def score(trace: list) -> dict:
    """Judge only the user-facing messages, ignoring tool chatter."""
    user_facing = [
        m["content"]
        for m in trace
        if m.get("role") in ("user", "assistant") and m.get("content")
    ]
    transcript = "\n\n".join(user_facing)

    result = tools.llm_judge(
        prompt="Was the assistant helpful in this conversation?\n\n{{ transcript }}",
        input={"transcript": transcript},
        model="gpt_4_1_mini",
        provider="openai",
    )
    return json.loads(result)
```

### Calling Other Tools

The LLM tools aren't special: a code judge can call any tool in its allowlist, using the same `from kiln import tools` API that [code tools](/docs/tools-and-mcp/code-tools) use. That lets a judge look up ground truth in your own systems, or reuse a tool you've already written.

See [Calling Other Tools](/docs/tools-and-mcp/code-tools#calling-other-tools) for the full API, including async calls and error handling.

### Timeout and Tools

* **Timeout**: wall-clock timeout for scoring one item, including any nested tool calls. Defaults to 180 seconds, with a 300 second maximum.
* **Tools**: an explicit allowlist of the tools your code may call. Code with no tools selected can't make tool calls at all.

### Your Code is a Real Python File

Kiln stores your judge's source as `scorer.py`, in the eval config's folder next to `eval_config.kiln`:

```
{task}/.../eval_configs/{id} - {name}/
  ├── eval_config.kiln   # metadata (no code)
  └── scorer.py          # your score() source
```

Because it's a plain Python file rather than a string inside JSON, it's importable, lintable, type-checkable, and produces a readable diff when you review changes in [Git](/docs/collaboration).

### Testing with pytest

You can write standard `pytest` tests against your judge. `score()` takes plain keyword arguments and depends only on the standard library plus `kiln_ai`, so you can import and call it directly — no Kiln-specific test runner.

Create `test_scorer.py` beside `scorer.py`:

```python
from scorer import score


# For an eval whose output score is named "Valid Output" (JSON key: valid_output).
def test_accepts_good_output():
    result = score(output='{"summary": "A real summary."}')
    assert result["valid_output"] == 1.0


def test_rejects_empty_summary():
    result = score(output='{"summary": ""}')
    assert result["valid_output"] == 0.0
```

Then run `pytest` from that folder. Kiln doesn't store, display, or run these tests — they live in a normal Python environment with `kiln_ai` installed.

### Learn More

* [Judge Types](/docs/evals-and-specs/judge-types): all the judge types, and when to use each
* [Code Tools](/docs/tools-and-mcp/code-tools): the same Python authoring model, for tools your agents call
* [Code Tools Authoring Guide](https://github.com/Kiln-AI/Kiln/blob/main/docs/code_tools_guide.md): the complete authoring contract, including the tool calling API and testing details


# Evaluate RAG Accuracy: Q\&A Evals

Know your Kiln Search tools find the right answer with RAG evals and synthetic Q\&A data

{% hint style="success" %}
**Evaluating RAG is tricky**. An LLM-as-judge doesn't have the knowledge from your documents, so it can't tell if a response is actually correct. But giving the judge access to RAG biases the evaluation.

**The solution** is reference-answer evals. The judge compares results to a known correct answer. Building these datasets used to be a long manual process, but Kiln makes the process fast and easy.
{% endhint %}

#### Overview

Reference answer accuracy evals measure how well your model leverages search tools (RAG) by generating query-answer (Q\&A) pairs from your document library and using them as reference answers to test your RAG system's responses. This approach includes:

* Generating large eval datasets quickly from your existing documents using synthetic Q\&A pair generation
* Creating realistic queries that reflect user questions about your corpus
* Using reference answers (ground truth) derived from your documents to evaluate accuracy
* Systematically testing different search tool configurations (chunking strategies, embedding models, etc.) and task models to find optimal settings

### Video Preview

{% embed url="<https://vimeo.com/1137040663>" %}

### The Workflow

This guide walks through the RAG-specific workflow for reference answer accuracy evals:

* Creating a Reference Answer Accuracy Eval
* Generating Q\&A pairs from your documents
* Setting up a Judge
* Finding the Ideal Run Configuration

For general eval concepts like judges, run configurations, and comparing results, see Evaluations.

<figure><img src="/files/8lsWfkHQ5wbZoZKjYB8a" alt="" width="375"><figcaption><p>The RAG Eval Process</p></figcaption></figure>

### Creating a Reference Answer Accuracy Eval

From the "Eval" tab in Kiln's UI, create a new evaluator using the "Reference Answer Accuracy Eval (RAG)" template.

{% hint style="info" %}
**Reference Answer Accuracy Eval (RAG)**: This template is designed for evaluating Q\&A pairs and includes a Reference Answer Accuracy score (pass/fail) that evaluates if the model's output is accurate as per the reference answer. The template is configured to work with Q\&A datasets built from your documents.
{% endhint %}

Select the template, edit if desired, and save your eval.

### Generate Q\&A Pairs

Most commonly, you'll want to populate your eval dataset using synthetic Q\&A pairs generated from your documents. These pairs include reference answers that serve as ground truth for evaluation. Clicking "Add Eval Data" from the Evals UI, and selecting "Synthetic Data" will launch the Q\&A generation tool with the proper eval tags already populated.

#### Select Documents

Choose which documents from your library to use for generating Q\&A pairs:

* **All documents**: Use every document in your library
* **Filter by tags**: Select specific documents by applying tag filters. This is useful when you want to generate evals for a specific subset of your corpus.

You can add tags to documents in the Document Library UI (found in the Docs & Search tab) to organize them for filtering.

#### Extract Documents

Before generating Q\&A pairs, you need to extract text from your documents. Choose an extractor config that will process your documents. The extractor converts your documents (PDFs, HTML, etc.) into markdown or plain text. If you've already extracted documents with this extractor, those extractions will be reused.

#### Generate Pairs

Configure the generation process to create Q\&A pairs from your documents.

**Generation Settings**

* **Pairs per document/chunk**: How many Q\&A pairs to generate from each document/chunk. More pairs give you a larger eval dataset but take longer to generate.
* **Model and provider**: The AI model used to generate Q\&A pairs. Larger models typically produce higher quality pairs.
* **Guidance**: Optional instructions to steer the generation. You can:
  * Use the default Q\&A generation template (recommended for most cases)
  * Provide custom guidance to focus on specific types of queries (e.g. "Focus on technical questions")

<details>

<summary>Optional: Break Long Documents Into <strong>Chunks</strong></summary>

Under "Advanced Options", check the "Split documents into smaller chunks" checkbox if you want to split documents before generating pairs. This is recommended for very long documents such as books, manuals, or transcripts. Splitting into smaller chunks helps create more focused Q\&A pairs.

* **Chunk size (tokens)**: The maximum size of each chunk. Smaller chunks create more focused Q\&A pairs but may miss context that spans chunks.
* **Chunk overlap (tokens)**: How much text overlaps between adjacent chunks. Overlap helps preserve context at chunk boundaries.

If you leave it unchecked, Q\&A pairs will be generated from entire documents without chunking. This works well for shorter documents or when you want queries that span the full document context.

**Generation Settings**

* **Pairs per document/chunk**: How many Q\&A pairs to generate from each document/chunk. More pairs give you a larger eval dataset but take longer to generate.
* **Model and provider**: The AI model used to generate Q\&A pairs. Larger models typically produce higher quality pairs.
* **Guidance**: Optional instructions to steer the generation. You can:
  * Use the default Q\&A generation template (recommended for most cases)
  * Provide custom guidance to focus on specific types of queries (e.g. "Focus on technical questions")

</details>

**What Gets Generated**

* **Queries**: Realistic questions that users might ask about your document corpus. These can be:
  * Natural language questions (e.g. "What is the population of Pittsburgh?")
  * Search-style queries (e.g., "Pittsburgh population 2020")
* **Reference Answers**: Factual, concise answers derived from the document content. These serve as ground truth for evaluating your RAG system's accuracy.

#### Review and Save

* Review generated pairs organized by document/chunk
* Remove individual pairs, entire chunks or documents if needed
* Click "Save All" to save the Q\&A pairs to your dataset.

Since these pairs contain reference answers, there's no need for a separate golden set. All pairs should be tagged with your eval tag (typically starting with `qna_eval_set`).

Pairs will also be saved with tags that identify:

* They're synthetic Q\&A data (`synthetic`, `qna`)
* Their generation session ID (starting with `synthetic_qna_session`)

### Setting up a Judge

Before evaluating different run configurations, you need to create a judge. The eval you created defines the goal, but the judge defines how it's run (judge type, and for an LLM judge the model and instructions).

Click "Create Judge" to get started. For detailed guidance on selecting a judge model and customizing evaluation instructions, see our [LLM Judges](/docs/evals-and-specs/llm-judges) guide. For the other judge types available, see [Judge Types](/docs/evals-and-specs/judge-types).

### Finding the Ideal Run Configuration

Once you have a judge set up, you can evaluate different configurations for running your RAG task. You can test different task models, prompts, and model parameters to find the best combination for answering questions from your document corpus. For detailed guidance on selecting and comparing task model options, see Finding the Ideal Run Method.

Since reference answer accuracy evals specifically test how well your RAG system retrieves and uses information from your documents, you'll also want to test different search tool configurations:

* A range of extraction models
* Different chunking strategies: fixed window or semantic, with varying chunk sizes and overlap
* A range of embedding models for search
* Different search index configurations: full-text search, vector search, or hybrid search, with varying K values
* A range of reranking models, with varying N top results

Once you've defined a set of run configurations (combining different task model options and search tool configurations), click "Run Eval" to test them against your Q\&A dataset.

<a href="/pages/Dtsf7EZfJgPB9VbSUiae#optimizing-your-rag" class="button primary">RAG Optimization Guide</a>

**Comparing Results**

After the eval completes, you'll see average scores for each run configuration. The highest average score indicates the best-performing configuration for your RAG system.

This systematic approach helps you find the optimal combination of task model and Search Tool settings for answering questions from your document corpus.


# Evaluate Appropriate Tool Use

Test your model's ability to appropriately invoke tools

<figure><img src="/files/ZfW2yft636OK3RBNDfSu" alt=""><figcaption></figcaption></figure>

{% hint style="success" %}
**It doesn't matter how well a tool works if it isn't invoked when needed.**

Kiln can evaluate your agent's tool use to help find issues early.
{% endhint %}

### **Overview**

Tool use evals measure how well your model decides when to invoke a tool. They check that:

* The tool is called when needed
* The tool isn’t called when it shouldn’t be
* The parameters passed to the tool are correct
* The agent doesn't take an unreasonable number of steps to get there

### Ways to Evaluate Tool Use

Most tool use questions have a definite right answer, and Kiln can check those directly in code:

* [**Tool Call Check**](/docs/evals-and-specs/judge-types#tool-call-check): inspects the agent's trace and asserts which tools were called, in what order, and with what arguments. No model call, so it's free, instant, and returns the same answer every time. **Start here.**
* [**Step Count Check**](/docs/evals-and-specs/judge-types#step-count-check): asserts the agent finished within a reasonable number of tool calls or turns. A good companion to a Tool Call Check — "it called the right tool" and "it didn't call it eleven times" are different bugs.
* [**Code Judge**](/docs/evals-and-specs/code-judges): for rules that are still deterministic, but go beyond what a Tool Call Check can express — comparing arguments across several calls, checking a tool's result rather than its arguments, or any logic you'd rather write in Python.
* [**LLM as Judge**](/docs/evals-and-specs/judge-types#llm-as-judge): a model reads the whole trace and forms an opinion. Reach for this only when whether the tool *should* have been called is genuinely a judgement call.

{% hint style="info" %}
The rule of thumb: if you can write down the tool calls you expect, use a Tool Call Check. If deciding whether the agent was right requires reading the user's request and thinking about it, use an LLM judge.
{% endhint %}

### Evaluating Tool Use with a Tool Call Check

#### Creating the Eval

From the "Evals" tab in Kiln's UI, create a new eval and select "Tool Call Check" under "Programmatic Judges".

You'll configure:

* **Expected Tools**: the tools you expect the agent to call. For each one, you can also specify expected arguments, matched exactly, by substring, or by regular expression.
* **Match Mode**:
  * All (any order): every expected tool was called
  * Any: at least one expected tool was called
  * Ordered (in list order): the expected tools were called, in the order you listed them
  * Never: none of the listed tools were called
* **Unlisted Tool Calls**: allow any other tool the agent calls, or fail the check if it calls anything not on your list.

Some examples of what these combinations express:

| Goal                                           | Configuration                                        |
| ---------------------------------------------- | ---------------------------------------------------- |
| The agent must search before answering         | `search`, match mode All                             |
| The agent must search, then fetch the page     | `search`, `fetch_url`, match mode Ordered            |
| The agent must never email a customer directly | `send_email`, match mode Never                       |
| The agent must stay within its allowed tools   | your tool list, with Unlisted Tool Calls set to fail |

#### No Golden Set Required

A Tool Call Check doesn't approximate human judgement, so there's no judge to align: you don't need a golden dataset, human ratings, or the [Compare Judges](/docs/evals-and-specs/evaluations#finding-the-ideal-judge) flow. Populate your test dataset and go straight to [comparing run methods](#finding-the-ideal-run-configuration).

### Evaluating Judgement with an LLM Judge

{% hint style="info" %}
LLM judges are more expensive and slower than Tool Call Check judges. Use them only when needed.
{% endhint %}

Some tool use questions really are subjective. "Should the agent have searched the web here, or was its own knowledge good enough?" doesn't have one right answer for every input — it depends on the request.

For those, create an eval with the "Desired Behaviour" template and add an [LLM as Judge](/docs/evals-and-specs/judge-types#llm-as-judge). Make sure conversation history is included, so the judge examines the full trace rather than only the final response.

#### Describe When the Tool Should be Used

Your judge needs guidelines describing correct usage. These help both the synthetic data generator create relevant test cases and the judge understand what "appropriate" means for your task.

**Web Research** tool examples:

* "Questions requiring current information, recent events, or up-to-date data that may have changed"
* "Questions about specific facts, statistics, or information that may not be in the model's training data"
* "Questions asking for comparisons, reviews, or analysis that benefit from multiple sources"

It's just as useful to describe when the tool should **not** be used. This creates a more comprehensive test dataset that includes clear negative cases.

**Web Research** tool examples:

* "General conversation, greetings, or questions that don't require factual information"
* "Questions about well-established facts, definitions, or concepts that are reliably in the model's training data"
* "Simple calculations, math problems, or questions that can be answered without external information"

#### Add Human Ratings

An LLM judge needs a golden dataset with human ratings, so you can measure how well it matches human preference. When rating for tool use, you'll need to check both the dataset entry and the Message Trace to see if the tool was invoked and with what parameters.

1. Click the `Rate Golden Dataset` button on the eval screen to go to the dataset view filtered to your golden dataset. For general guidance on rating, see [Reviewing and Rating](/docs/reviewing-and-rating).
2. **View the Message Trace** for each dataset entry to inspect:
   * Whether the tool was called
   * What parameters were passed to the tool

This allows you to accurately rate whether the tool was called appropriately based on the full conversation context.

For your eval output rating, you should click "Pass" if the model's behaviour was appropriate and "Fail" if not.

**Pass = Appropriate Tool Use**

* The model called the tool with correct parameters at the appropriate time
* The model correctly did not call the tool as it was not needed

**Fail = Inappropriate Tool Use**

* The model should have called the tool but did not
* The model called the tool but shouldn't have, or called it with wrong parameters.

#### Finding the Ideal Judge

For detailed guidance on selecting judge models and customizing evaluation steps, see our [LLM Judges](/docs/evals-and-specs/llm-judges) guide. For comparing judges to find the one that best aligns with human preferences, see [Finding the Ideal Judge](/docs/evals-and-specs/evaluations#finding-the-ideal-judge).

{% hint style="success" %}
You can add both to one eval. A Tool Call Check that asserts the mechanics, plus an LLM judge for the judgement call, gives you a cheap signal that runs on everything and a nuanced one where it matters.
{% endhint %}

### Generate Synthetic Test Data

Populate your test dataset using synthetic data generation. Click "Add Eval Data" from the Evals UI and select "Synthetic Data" to launch the generator with the proper eval tags already populated.

{% hint style="info" %}
**Tool-Specific Behaviour**: when generating outputs, the tools available to your task are enabled for this step, ensuring your model has the opportunity to call them when appropriate. The system captures the full conversation trace, including whether the tool was called and what parameters were used.

See our [Synthetic Data Generation docs](/docs/synthetic-data-generation) for more guidance.
{% endhint %}

### Finding the Ideal Run Configuration

Once your eval is set up, you can evaluate different configurations for running your task. Since tool use evals specifically test your model's tool calling behaviour, you'll want to test configurations such as:

* Different task models (some models are better at tool calling than others)
* Different prompts and system instructions
* Different sets of available tools (testing with varying numbers of tools can help identify optimal tool configurations)

You will also be able to see the tools available to each run configuration. Keep in mind that any run configuration without access to the tool you are evaluating will produce unreliable scores, as the model cannot actually invoke the tool being tested.

For detailed guidance on selecting and comparing task model options, see [Finding the Ideal Run Method](/docs/evals-and-specs/evaluations#finding-the-ideal-run-method).


# Tools & MCP

Connect powerful tools to your Kiln tasks

Kiln allows connecting to tools such as [Kiln Search Tools (RAG)](/docs/documents-and-search-rag) or third party tools via via [Model Context Protocol (MCP)](https://modelcontextprotocol.io). You can also write your own tools in Python with [Code Tools](/docs/tools-and-mcp/code-tools). These tools can give your Kiln tasks powerful new capabilities.

## Video Walkthrough

{% embed url="<https://www.youtube.com/watch?v=qh0FIrLMrII>" %}

## Connecting Tools

First, connect some tools to your Kiln project. Tools are connected at the project level, and become available to all tasks in the project.

To connect a new set of tools, open "Tools" > "Add Tools".

### Quickstart: Math Tools

If you want to try tools as quickly as possible, enable the "Math Tools" from the "Add Tools" screen. These built into Kiln and don't require setting up a MCP server. They can be enabled in one click, and will add 4 simple math tools to your project: add, subtract, multiply and divide.

### Kiln Search Tools (RAG)

Kiln allows you to build powerful search tools, which can search thousands of documents for knowledge. These behave like any other tool: each search tool you create, will be automatically available in the tools dropdown.

See the [Documents & Search](/docs/documents-and-search-rag) docs for details.

### Code Tools

You can write your own tools in Python, without leaving Kiln. A code tool can call other tools it's been granted access to, which makes it ideal for batching, filtering, and cleaning up messy interfaces before the model ever sees them.

Open "Tools" > "Add Tools" > "Code Tool" to create one. See the [Code Tools](/docs/tools-and-mcp/code-tools) docs for details.

### Powerful Example Tools: Web Search, Python Interpreter, and more

{% hint style="info" %}
The example MCP servers are not part of Kiln. You should trust the authors, their code, and their privacy policies before enabling them.
{% endhint %}

Kiln has several popular MCP servers pre-configured. You can get started with them in just a few clicks:

* [Web Search and Scrape by Firecrawl](https://docs.firecrawl.dev/mcp-server): Allows your task to search the web and scrape websites into text. Requires an API key from Firecrawl.
* [Run Python Code by Pydantic](https://ai.pydantic.dev/mcp/run-python/): Allows your model to generate then run python code. Powerful for allowing LLMs to perform tasks they normally don’t excel at: complex math, iterative computation, matrix math, etc. The python interpreter is sandboxed, and can’t access your system.
* [Local File Access by Anthropic](https://github.com/modelcontextprotocol/servers/tree/HEAD/src/filesystem): Allow your tasks to access files in specific folders on your local hard drive. You’ll be able to specify which folders it can access during setup.
* [Stock Quotes by Twelve Data](https://github.com/twelvedata/mcp): Realtime access to stock quotes and other market data. Requires a Twelve Data API key.
* [Control Github by Github](https://github.com/github/github-mcp-server): Manage repos, issues, PRs and workflows. Requires a GitHub Access Token.

<figure><img src="/files/zl6D9Af6WRZdBNEhKNSu" alt=""><figcaption><p>Example tools on the "Add Tools" screen</p></figcaption></figure>

You may need to install developer tools like Node.js or Deno before running these MCP servers. In each case, there’s a header at the top with a link to the appropriate dependency:

<figure><img src="/files/HHSU3ryM24i2ey5s97uC" alt="" width="375"><figcaption><p>Example dependency warning</p></figcaption></figure>

### Adding Custom MCP Servers

You can connect any MCP server to Kiln! MCP servers come in two flavours:

* Local servers: run locally on your computer, and are defined by a terminal command and arguments.
* Remote servers: run on a server and you connect to them over the web/http.

You can discover new MCP servers to use in registries like [MCP Pulse (external)](https://www.pulsemcp.com/servers).

#### Connecting Local MCP Servers

Once you’ve found a local MCP server you want to connect, click “Tools” > “Add Tools” > “Local MCP” > “Connect”, and provide the appropriate information in the setup:

* Name and Description: fields for you and your team to identify the server
* Command: The command to run. Just the actual command, not including arguments. For example: `npx`, `uvx`, `deno`, or similar. It should not include spaces.
* Arguments: any arguments to pass to the command, for example the `-y firecrawl-mcp` portion of the command `npx -y firecrawl-mcp`
* Environment variables: a list of environment variables. Be sure to specify which are secrets (like API keys), so that these aren't synced into your Kiln project on Git ([more info](#secret-management)).

{% hint style="info" %}
Be sure the server runs in [stdio mode](https://modelcontextprotocol.io/specification/2025-06-18/basic/transports#stdio). This is normally the default, but mode is sometimes is exposed as an argument.
{% endhint %}

<details>

<summary>Advanced Troubleshooting</summary>

Kiln will attempt to use your standard PATH to find the appropriate commands. If the command works in a fresh terminal window, they should work in Kiln. If you’re having issues:

1. Ensure the same command works in a new terminal window. If not, debug that by installing any missing dependencies, or adding the required commands to your PATH.
2. Ensure the server is running in [stdio mode](https://modelcontextprotocol.io/specification/2025-06-18/basic/transports#stdio), following its documentation.
3. If you need to specify a custom PATH variable for Kiln, you can set the custom\_mcp\_path variable in the YAML file \~/.kiln\_ai/settings.yaml to set PATH for all Kiln projects, or set PATH manually as an environment variable in a specific local MCP server config.

</details>

#### Connecting Remote MCP Servers

Once you’ve found a remote MCP server you want to connect, click "Tools” > “Add Tools” > “Remote MCP” > “Connect”, and provide the appropriate information in the setup:

* Name and Description: fields for you and your team to identify the server
* Server URL: The URL to the remote MCP server, for example <https://api.githubcopilot.com/mcp/>
* Headers: a list of headers to pass when calling the remote MCP server. Be sure to specify which are secrets (like API keys), so that these aren't synced into your Kiln project on Git ([more info](#secret-management)).

## Using Tools

Once you’ve connected tools, you can use them in any of your tasks. On the “Run” screen you can select which tools will be available to the model under “Advanced”. You can select any number of tools, then run your task as you normally would.

<figure><img src="/files/25pZG70r9GAavrMXf6xD" alt="" width="182"><figcaption></figcaption></figure>

### Viewing Tool Calls

{% hint style="info" %}
The model may or may not choose to use the tools you provide. If you find the model is not using tools when you feel it should, update your prompt to more explicitly specify when tools should/must be used.
{% endhint %}

Tool calls happen behind the scenes, between the user message and the model's final output.

If you want to view which tools were called, their arguments, and their results look at the "All Messages" section on the "Run" screen or when viewing a dataset entry. You'll be able to see tool calls the model makes, as well as the tool's response:

<figure><img src="/files/4sXsmUl0WAU75Wb26bYG" alt="A Trace Including Tool Calls" width="375"><figcaption><p>A Trace Showing a Tool Call</p></figcaption></figure>

## Guidance / Understanding Tools

### Always Select Models Optimized for Tool Calling

Not all models are designed with tools in mind. Calling tools requires precise syntax and models that were not trained for tool-calling often fail. This makes it important to select a model with tool support.

Kiln tests every model for its capabilities to minimize guesswork; you can read more about our testing methodology [here](https://kiln.tech/blog/i_wrote_2000_llm_test_cases_so_you_dont_have_to). You can view our [model library](mailto:undefined) and filter to Capability=Tools to see which models and providers we suggest for tool calling.

When you select a model in Kiln, you’ll see a warning if tools are not supported. While you can still try these models, they are likely to fail:

<figure><img src="/files/A6S1zjQiMO8JRmSt58tB" alt="" width="375"><figcaption><p>Ignore the warnings at your peril</p></figcaption></figure>

### Don’t Add Too Many Tools

When you add tools to a task run, the model processes their names and descriptions as part of its context. Adding a few relevant tools can be a great way to improve task performance, but if you add dozens or hundreds you’ll both fill your context and dilute the model’s attention. The right number will depend on your task and the base model, but be selective when adding tools.

See our [video guide at 7m23s](https://youtu.be/qh0FIrLMrII?si=CHiJ1CCen_2MA5A3\&t=443) for a discussion of these concepts:

<figure><img src="/files/4H57OxBaUC0CI6hbANL2" alt=""><figcaption></figcaption></figure>

### Synthetic Data & Fine-Tuning with Tools

At this time, Kiln doesn’t support tool calling for synthetic data and fine-tuning. Rest assured, we’re working on adding it!

### Secret Management

Often MCP servers require secrets, like API keys. However, Kiln projects are designed to be [shared across teams with Git](/docs/collaboration) and you don’t want to commit secrets in a Git repo.

To address this, Kiln allows you to mark headers and environment variables as secrets. Secrets are never stored in the Kiln project files, which are designed to be shared/synced. If a team member adds a tool which requires secrets, anyone who syncs it will have to re-enter the secrets in settings before using the tool. Non-secret headers and environment variables will be synced automatically.

<figure><img src="/files/5J99dT3EynWLxX4TCQZB" alt="" width="375"><figcaption><p>Identify secrets on tool setup</p></figcaption></figure>

## Kiln MCP Server

The documentation above is for using Kiln as a MCP client (where Kiln calls other MCP servers to use tools). Kiln can also function as a MCP server (where Kiln exposes tools to other MCP clients).

Kiln can expose tools for:

* Any [Search Tools (RAG)](/docs/documents-and-search-rag) you have created in Kiln
* Any [Kiln Task as Tool](/docs/agents#multi-actor-interaction-aka-subtasks) agents you have created in Kiln

To use Kiln as a tool server

* Install Kiln CLI server tools: `uv tool install kiln_server`
* Run the `kiln_mcp` command, pointing it to the Kiln project you want to host. See [the docs on the Kiln MCP Server](https://github.com/Kiln-AI/Kiln/tree/main/libs/server/kiln_server/mcp#readme) for details.


# Code Tools

Write a Python function that runs as a tool, and can call other tools

Code tools let you write Python that runs as a tool inside your agent, without leaving Kiln. They're stored in your project like any other artifact, and appear in the tools dropdown alongside your MCP servers and search tools.

### Why Code Tools

Agents rarely fail because the model is bad at reasoning. They fail because the interface they're given is messy:

* **N+1 tool loops**: an API with no batch endpoint forces the agent to call `get_user` fifty times, one per turn, burning context and time.
* **Context floods**: a tool returns a 40KB JSON blob when the agent needed three fields from it.
* **Missing operations**: the thing the agent actually needs is two API calls and a join, and no single tool does it.

You can try to fix these in the prompt, and the model will get it right most of the time. A code tool fixes them in code, once, and gets it right every time. A batch wrapper, a result filter, or a purpose-built endpoint is a durable artifact — it doesn't drift when you change models.

{% hint style="info" %}
Code tools run inside the Kiln desktop app, on your machine. They aren't available to server-side deployments.
{% endhint %}

### Creating a Code Tool

Open "Tools" > "Add Tools" > "Code Tool", and Kiln will walk you through two steps.

#### Step 1: Define the Tool

* **Display Name**: the name you and your team see in Kiln. For example "User Lookup".
* **Tool Name**: the function name exposed to the model. Lowercase with underscores, for example `get_user`.
* **Description**: shown to the model. Describe what the tool does and when to use it — this is what the model reads when deciding whether to call it.
* **Parameters**: the arguments the model passes to your tool, defined with the same schema builder used for task inputs and outputs.

{% hint style="warning" %}
**Write the description for the model, not for yourself.**

Like [skills](/docs/skills#creating-skills), the description is the model's only signal about when to use this tool. "Look up a user by ID and return their profile" beats "user tool".
{% endhint %}

#### Step 2: Write the Code

Write a function named `run`. Your parameters arrive as keyword arguments matching the schema you defined:

```python
def run(user_id: str) -> str:
    """Look up a user and return their profile."""
    return f"User {user_id}"
```

Both sync and async forms work — use `async def run(...)` when you want concurrency.

Return values become the tool output the model sees. Strings pass through as-is; dicts, lists, numbers and booleans are JSON-serialized for you.

The right-hand panel of the editor has two things worth using before you save:

* **Tool Access**: an allowlist of the tools your code may call. Code with nothing selected can't make tool calls at all.
* **Test panel**: run your tool against real arguments, calling your real allowlisted tools, and see what comes back.

Under "Advanced Options" you can set the **timeout**: the wall-clock limit for one invocation, including any nested tool calls. Defaults to 60 seconds.

### Calling Other Tools

The most powerful thing a code tool can do is call other tools. Two modules are available in your code:

```python
from kiln import tools          # sync — blocks until the tool returns
from kiln import async_tools    # async — awaitable, concurrent under gather
```

Call an allowlisted tool as an attribute, with keyword arguments:

```python
import json
from kiln import tools


def run(query: str, max_results: int = 10) -> str:
    """Search, then return only the fields the agent actually needs."""
    raw = tools.search(query=query)
    results = json.loads(raw)

    filtered = [
        {"title": r["title"], "url": r["url"]}
        for r in results[:max_results]
        if "title" in r and "url" in r
    ]
    return json.dumps(filtered)
```

A few rules worth knowing up front:

* **Tool calls always return a string** — byte for byte what the model would have seen. Parse it yourself with `json.loads` when the tool returns JSON.
* **Keyword arguments only.** Positional arguments raise an error.
* **Only allowlisted tools resolve.** Calling anything else raises `ToolNotAllowed`, and the error lists what is available.
* **`tools.list_tools()`** returns the tools in your allowlist, with their descriptions and parameter schemas.

For true concurrency, use `async_tools` with `asyncio.gather` — this is the fix for the N+1 loop:

```python
import json
import asyncio
from kiln import async_tools


async def run(user_ids: list[str]) -> str:
    """Fetch many users at once, instead of one agent turn per user."""
    async def fetch(uid):
        return json.loads(await async_tools.get_user(id=uid))

    users = await asyncio.gather(*(fetch(uid) for uid in user_ids))
    return json.dumps(users)
```

Tool calls raise typed exceptions you can catch for retries: `ToolNotAllowed`, `ToolTimeout`, and `ToolCallError`. Import them from `kiln.tools` or `kiln.async_tools`.

{% hint style="info" %}
Code tools can call other code tools — they're just tools.
{% endhint %}

### Code Trust

{% hint style="danger" %}
**Code tools run on your machine with full access.**

There is no sandbox: no import restrictions and no resource limits beyond the wall-clock timeout. Code can read your files, make network calls, and anything else Python can do.
{% endhint %}

The trust gate is the security boundary. Adding or editing code in a project requires confirming that you trust it; running code you've already saved and trusted doesn't prompt you again. Trust is granted for the session, so Kiln asks again after a restart.

This matters most when a project came from somewhere else. Kiln projects are designed to be [shared across teams with Git](/docs/collaboration), and a project you sync or import may contain code tools written by someone else. Importing a project authored elsewhere requires explicit approval before its code will run — treat that prompt the way you'd treat running an unfamiliar script from the internet, and read the code first.

### Your Code is a Real Python File

Kiln stores your tool's source as `tool.py`, in the tool's folder next to `code_tool.kiln`:

```
{project}/code_tools/{id} - {name}/
  ├── code_tool.kiln   # metadata (no code)
  └── tool.py          # your source — byte-for-byte what runs
```

Because it's a plain Python file rather than a string inside JSON, it's importable, lintable, type-checkable, and produces a readable diff in Git.

### Testing with pytest

Kiln ships a `pytest` plugin, so you can write standard tests against your tool in a normal Python environment. Install `kiln_ai` (`pip install kiln-ai`) and the plugin is auto-discovered: the `from kiln import tools` at the top of your `tool.py` resolves under `pytest`, and a `kiln_tools` fixture becomes available for stubbing tool responses.

Create `test_tool.py` in the same folder as `tool.py`:

```python
import json
import tool  # the artifact's tool.py — imports cleanly under pytest


def test_filters_results(kiln_tools):
    kiln_tools.set("search", json.dumps([
        {"title": "Real", "url": "https://example.com"},
        {"title": "No URL"},
    ]))

    out = json.loads(tool.run(query="anything"))

    assert out == [{"title": "Real", "url": "https://example.com"}]
    assert kiln_tools.calls[0].name == "search"
```

The fixture stubs replies with `kiln_tools.set(name, reply)`, forces errors with `kiln_tools.set_error(name, exc)`, and records every call in `kiln_tools.calls`. It behaves like the real runtime: unregistered tools raise `ToolNotAllowed`, positional arguments raise `ToolCallError`.

Kiln doesn't store, display, or run these tests — the loop lives in your own Python environment.

{% hint style="info" %}
`tool.py` is a fixed filename, so running `pytest` across several tool folders at once hits a module name collision. Run `pytest` from inside a single tool folder, or use `pytest --import-mode=importlib`.
{% endhint %}

### Code Tools vs Other Options

|                   | Code Tool                                       | MCP Server                                                | Kiln Task as Tool                         |
| ----------------- | ----------------------------------------------- | --------------------------------------------------------- | ----------------------------------------- |
| **Best for**      | Wrapping, batching and filtering existing tools | Connecting an existing service or third-party integration | Delegating a sub-problem to another agent |
| **Written in**    | Python, inside Kiln                             | Any language, outside Kiln                                | A Kiln task prompt                        |
| **Deterministic** | Yes                                             | Depends on the server                                     | No — it's a model call                    |
| **Setup**         | Low                                             | Medium                                                    | Medium                                    |

### Learn More

* [Code Tools Authoring Guide](https://github.com/Kiln-AI/Kiln/blob/main/docs/code_tools_guide.md): the complete authoring contract, with more examples of concurrency, retries and error handling
* [Code Judges](/docs/evals-and-specs/code-judges): the same Python model, used to score evals


# Connect to Existing Agents

You can connect Kiln Evals to existing agents on any platform, using MCP

If you've already built an agent using another platform, you can still use Kiln to evaluate and optimize your system!

### Step 1: Expose your Existing Agent as an MCP Server

Wrap your existing agent in an [MCP server](https://modelcontextprotocol.io/docs/getting-started/intro), so Kiln can call it. You can use any library, framework, or programming language.

Exactly how to do this will depend on how you implemented your existing agents, but every major programming language has an MCP library for wrapping existing functions as MCP servers:

* [Python](https://github.com/modelcontextprotocol/python-sdk?utm_source=chatgpt.com)
* [Javascript/Typescript](https://github.com/modelcontextprotocol/typescript-sdk?utm_source=chatgpt.com)
* [Go](https://github.com/modelcontextprotocol/go-sdk)
* [Rust](https://github.com/modelcontextprotocol/rust-sdk?utm_source=chatgpt.com)
* [Java](https://github.com/modelcontextprotocol/java-sdk)
* [More](https://github.com/modelcontextprotocol)

### Step 2: Connect Your MCP Server to Kiln

Follow [our instructions](/docs/tools-and-mcp#connecting-tools) to connect your MCP server to Kiln.

### Step 3: Connect an MCP Tool to a Kiln Task

A "task" in Kiln represents a unit of work an agent can do. It has a specific input/output schema, and may have a collection of evals to measure quality. There are a number of ways to connect an MCP tool to your Kiln tasks.

#### Option 1: Create a New Kiln Task From an MCP Tool

The easiest way is to create a new task from your MCP tool signature. It will create a new task with the exact same input and output schema as your MCP tool.

This is the preferred option if you're just getting started and don't already have evals created for an existing task. However, if you already have evals and Kiln agents, you'll probably want to use Option 2 or Option 3.

To create a new task from a tool, click `Settings > Manage Tools > Tool Server > Individual Tool > [...] Menu > Create Task from Tool`.

<figure><img src="/files/5bRjkOhDya0dPckB8z0p" alt="" width="276"><figcaption></figcaption></figure>

#### Option 2: Use a Wrapper Agent

Any existing Kiln task can call your MCP tools. You can give your existing agents in Kiln access to your external MCP tools by selecting them in the `Tools & Search` dropdown on the run screen.

<figure><img src="/files/YgKkCNw6JjUnSWUVE3Ad" alt="Select Tools on the Run Screen" width="375"><figcaption></figcaption></figure>

This is a great approach if your task's input/output schemas doesn't exactly match your tool's input/output schemas. The wrapping LLM agent can follow instructions, and call the tool with the needed parameters.

#### Option 3: Connect an MCP Tool to an Existing Task

If you have an existing Kiln task and want to use Kiln to compare multiple methods of running your task (MCP tools, Kiln agents), you can connect your MCP tool to an existing task as a "Run Configuration". Run configurations describe how to run a specific task. For an agent, it's typically fields like model, system prompt, and LLM parameters; however, a run configuration can also point to any MCP server.

{% hint style="info" %}
The input/output signature of your MCP tool **must exactly match** the input/output signature of your Kiln task, or you won't be able to directly connect a task to an MCP tool.

If it doesn't match, we suggest either using a wrapper agent (option 2) or modifying the MCP tool's schema to align to your task.
{% endhint %}

To connect an MCP tool to an existing task, click `Settings > Manage Tools > Tool Server > Individual Tool > [...] Menu > Run Task with Tool > Run Task Directly`.

<figure><img src="/files/By6Kv7KC1QOpKOS4NFrY" alt="" width="375"><figcaption></figcaption></figure>


# Prompts

How to use our prompt optimizer, prompt generators, or create your own prompt

Kiln offers several methods to build and manage prompts

* [Automatic Prompt Generator](/docs/prompts/automatic-prompt-optimizer): Our state-of-the-art automatic prompt optimizer. We run iterative experiments and use evals to pinpoint and fix failure modes — **no manual prompting required**.
* [Prompt Generators](#prompt-generators): Kiln can automatically generate many popular prompt styles from your task and dataset (few-shot, many-shot, chain of thought, chain of thought multi-shot, and more). The more you use your task, and rate the results, the richer your prompts become.
* [Custom Prompts](#custom-prompts): manually create, save and share any prompt.

<figure><img src="/files/CbupxmBr3WflBPKV5UxL" alt=""><figcaption></figcaption></figure>

## Viewing, Managing & Sharing Prompts

The "Prompts" tab in the UI lets you see and manage all of the prompts for the currently selected task.

* Create a new prompt
* View saved prompts
* Manage prompts (rename, delete)
* View prompt generators

Anyone you [collaborate](/docs/collaboration) with will automatically have access to the same set of prompts.

## Prompt Fields

When creating a new prompt, there are several fields:

* Name: a name for you and your team to identify this prompt. Not used by the model.
* Prompt (aka System Message): The core of your prompt. Will be passed to the model as a system message before any user data is sent.
* Chain of thought instructions: if provided, using this prompt will add an extra "thinking"/reasoning phase to its execution. These instructions guide how the model should "think" about the problem before answering. See [docs on Chain of Thought and Thinking](/docs/reasoning-and-chain-of-thought#chain-of-thought-call-flow-non-reasoning-model) for details.

## Using Prompts

You can select any available prompt or prompt generator from the prompt dropdown:

<figure><img src="/files/8tBuqfX4xg1RtT8zkhcH" alt="" width="310"><figcaption><p>Select a Prompt</p></figcaption></figure>

{% hint style="info" %}
You can start typing to filter this list, which can make it easy to find a prompt by name.
{% endhint %}

## When to Use Each Type of Prompt

Ultimately it's up to you when to use each style. The best approach varies from task to task, and model to model.

It's worth evaluating a range of prompt/model pairs to find one that works best for your task, while considering speed/cost tradeoffs of longer prompts and larger models. [Kiln Evals](/docs/evals-and-specs) give you a scientific way to find the best prompt for your task.

{% hint style="success" %}
**LLMs are often better at writing prompts than humans.** Given a good evaluator, they can test hundreds of unique prompts on thousands of test cased to find the ideal prompt.

**Try our** [**automatic prompt optimizer**](/docs/prompts/automatic-prompt-optimizer) to find the best prompt for your task, using evals.
{% endhint %}

#### Don't Assume the Same Prompt Will Work on Every Model

Different models will interpret a prompt differently. You may need to re-optimize your prompt when changing or upgrading your model.

#### Prompts for Fine Tuning

If fine-tuning we generally suggest a tiered approach:

* For generating training data: use a long/powerful prompt like "Chain of Though - Few Shot", on a powerful model (GPT, Claude, Deepseek).
* When building fine-tunes, try a range of included prompts, including the original prompt used when generating training data, the "Basic (Zero Shot)", and an even shorter custom fine-tune prompt. Also include a range of models and model sizes in your search (llama 1b, 3b, 8b, 70b, etc).
* Evaluate the resulting models. See if the longer prompts are necessary. It's possible the very short prompts will perform well after fine-tuning, which improves speed and lowers costs.
* Read more guidance from [OpenAI](https://platform.openai.com/docs/guides/fine-tuning#crafting-prompts)


# Automatic Prompt Optimizer

Find the best prompt for your task, automatically.

{% hint style="warning" %}
**Kiln Prompt Optimizer Requires a Kiln Pro Enterprise Plan**

The Kiln Prompt Optimizer runs on Kiln's servers, and consumes millions of tokens each run. Due to the high cost of running the optimizer, the prompt optimizer is a paid feature.

[Contact us](mailto:support@kiln.tech) to discuss an enterprise plan.
{% endhint %}

Kiln’s Prompt Optimizer automatically finds high-performing prompts for your task. It often beats manual prompt engineering by double-digit gains on evals.

### How It Works

To find the optimal prompt, Kiln Prompt Optimizer combines [Kiln Evals](/docs/evals-and-specs), synthetic training dataset generation, and algorithmic reflective prompt evolution.

Instead of human trial-and-error, Kiln will run thousands of automated experiments and iteratively find an optimal prompt for a given model and task.

#### Kiln Evals Drive Quality

We can't optimize something unless we can measure it, so the heart of our prompt optimizer is [Kiln Evals](/docs/evals-and-specs).

Follow our guides to create evals that measure your task's quality. The better your evals are at assessing quality, the better the prompt optimizer will work. Some guidance:

* [Create many small evals](https://kiln.tech/blog/you_need_many_small_evals_for_ai_products): it's typically easier to create several small evals focused on one area, than to try to create an all-encompassing eval
* Use the [Kiln Eval Builder](/docs/evals-and-specs/specifications) to make better evals: it uses AI to refine your LLM-as-Judge for better alignment to human preference

#### Synthetic Training Data and Withheld Eval Data

When you create an eval with the Eval Builder, we generate two separate datasets: **training** and **eval**. During optimization, we keep eval data withheld from training so results stay unbiased.

{% hint style="info" %}
**Legacy Evals May be Missing Training Data**

If you have an eval created before the Eval Builder was added, it may not have a training dataset. You'll need to use [synthetic data generation](/docs/synthetic-data-generation) to generate a training dataset, [tag the results](/docs/organizing-datasets#using-tags-to-organize-your-dataset), and save the tag as your eval's training dataset tag.
{% endhint %}

#### Reflective Prompt Evolution

Kiln’s Prompt Optimizer is powered by reflective prompt evolution (inspired by [GEPA](https://arxiv.org/abs/2507.19457), with additional optimizations).

At a high level, we start from your prompt as a baseline, then run **hundreds of iterative prompt mutations**. Each iteration is scored using evals to verify that changes improve performance, and to catch regressions early.

The process is conceptually similar to fine-tuning, but instead of updating model weights, it focuses entirely on improving the prompt.

### Guide

1. **Create one or more** [**evals**](/docs/evals-and-specs) that define the desired behaviour of your product.
2. **Choose a base model** that performs reasonably well on your task. If it can perform well on a naive prompt, it's more likely to improve with prompt optimization. The prompt we produce will be optimized for this specific model, and may not work as well on other models.
3. **Run the Prompt Optimizer** to evolve and validate improved prompts. Simply select "Create Prompt" in the UI and follow the steps.

### **Fine Tuning vs Prompt Optimization**

Prompt optimization is typically **faster and easier** than fine-tuning. We generally recommend optimizing your prompt first, because it:

* Produces strong results quickly
* Requires no hyperparameter tuning or data-science skills
* Makes overfitting easier to avoid and detect
* Is easier to deploy (just update your prompt)

|                         | Fine Tuning               | Prompt Optimization          |
| ----------------------- | ------------------------- | ---------------------------- |
| **Effort**              | High                      | Low                          |
| **Optimization Target** | Supervised Training Data  | Evals                        |
| **Time**                | 20m to 1 day              | \~1 hour                     |
| **Interpretability**    | Can't interpret changes   | Easy: read your new prompt   |
| **Deployment Effort**   | High: host a custom model | Low: just change your prompt |


# Prompt Generators

Prompt generators build prompts for you combing best practices and data from your dataset

<figure><img src="/files/CbupxmBr3WflBPKV5UxL" alt=""><figcaption></figcaption></figure>

### Prompt Generator Inputs

When you create a Kiln task you specify several fields which are used when generating prompts:

* Instructions: The core of your prompt. Describe the task as you would in a standard LLM prompt.
* Thinking Instructions: This field allows you to specify a portion of the prompt, which is only used when you select a "chain of thought" prompt template. If omitted, the default is "Think step by step, explaining your reasoning.".

These can be edited in `Settings > Edit Task > Task Instructions` .

### Adding Content for Multi-Shot Prompts

Multi-shot prompting is when you provide several examples in the prompt, and has been shown to greatly improve output quality.

In order to use multi-shot prompting in Kiln, we need some correct example content to use. Kiln makes this easy! Simply run your task in the Run tab of the UI, and rate the output. As you collect more 4 and 5 star responses, the set of available content for multi-shot prompting increases.

Kiln will automatically select content for multi-shot prompts using your dataset, using the following priority:

* 5-star [repairs](/docs/repairing-responses): where a human has given feedback, repaired the response, and rated the repair 5-stars
* 5-star responses
* 4-star response

Responses with a 3-star or worse rating are never used, unless they have been [repaired](/docs/repairing-responses) to 5-stars.

### Prompt Generators

Once you've setup your task and added some example content, you'll have a number of prompt generators to choose from:

* **Basic (Zero Shot)**: a prompt template that will only use your task definition (instructions and requirements). This prompt is deterministic, and won't change unless you edit your task.
* **Few Shot**: a multi-shot prompt that will include the 4 best examples from your dataset, on top of the basic prompt.
* **Many Shot**: a multi-shot prompt that will include the 25 best examples from your dataset, on top of the basic prompt.
* **Repair Multi Shot**: a multi-shot prompt that will include the 25 best examples from your dataset, on top of the basic prompt. This [prompt will use repaired examples](/docs/repairing-responses) to show 1) the generated content which had issues, 2) the human feedback about what was incorrect, 3) the corrected 5-star content. This gives the LLM examples of common errors to avoid.
* **Basic Chain of Thought**: a [chain of thought](/docs/reasoning-and-chain-of-thought) prompt template that will only use your task definition (instructions, requirements, and thinking instructions). Kiln will return the response in 2 stages, 1) thinking stage, 2) response stage. Both are stored in the dataset, but typically only stage 2 is used in the app/product. For structured tasks, the final answer must conform to the schema, but the thinking stage is plain text. This prompt is deterministic, and won't change unless you edit your task.
* **Chain of Thought - Few Shot**: a prompt that applies both the few shot template (4 examples), and chain-of-thought thinking instructions.
* **Chain of Thought - Many Shot**: a prompt that applies both the many shot template (25 examples), and chain-of-thought thinking instructions.


# Documents & Search (RAG)

Add knowledge to your AI systems with docs & search

RAG (Retrieval-Augmented Generation) is a powerful technique for adding knowledge to AI systems, and Kiln makes building RAG systems incredibly easy!

### Quick Start: Create a RAG in Under 5 Minutes

Building a search tool in Kiln takes 3 steps:

1. [Adding Documents](#adding-documents): Drag and drop files into the Kiln document library
2. [Create a search tool](#building-a-search-tool): specify how you want Kiln to search the documents
3. [Use the tool](#using-search-tools): Select the search tool when running your task

See our Video Guide if you prefer video format:

{% embed url="<https://vimeo.com/1138141451>" %}

### Document Library

You can open the document library from the “Manage Documents” link in the “Docs & Search” tab.

#### Adding Documents

To add documents, simply click “Add Documents” in the Document Library, then drag-and-drop in as many files as you like.

<figure><img src="/files/t7ODvWfje5Syydn8Vi0R" alt="" width="375"><figcaption><p>The Add Documents dialog</p></figcaption></figure>

{% hint style="info" %}
Documents are added to your Kiln project; if you’re using Kiln to [collaborate with a team](/docs/collaboration), documents will be available to everyone.
{% endhint %}

#### Supported File Types

Kiln supports the following file types:

* **Documents**: .pdf, .txt, .md, .html
* **Images**: .jpg, .jpeg, .png
* **Videos**: .mp4, .mov
* **Audio**: .mp3, .wav, .ogg

{% hint style="info" %}
Not every extraction model can handle every file type. Use Google Gemini models for maximal file type support. When creating a custom search tool, the model selection dropdown will list supported file types.
{% endhint %}

#### Tagging Documents

Documents can be organized by adding tags. This is typically used to sub-divide your docs library into sections, which allows you to build search tools targeting specific document sets. Here are some examples:

* **knowledge\_base**: your public help docs / knowledge base
* **customer\_support\_policies**: internal docs for how to respond to various types of customer requests
* **product\_specs**: feature definitions, product requirement docs, spec sheets
* **blog\_posts**: your company’s blog posts

You can add or remove tags in Kiln in 2 ways:

1. Single document: open a document’s page in the document library, then add or remove tags using the “Tags” sidebar.
2. Many documents: click “select” in the document library, select all relevant documents, click the tag icon, then select “Add Tags” or “Remove Tags”

<figure><img src="/files/VrLUPj8nEtO1WVzUU1GI" alt="" width="335"><figcaption><p>Managing tags from the document's detail page</p></figcaption></figure>

### Building a Search Tool

Once you’ve added documents, you can create a search tool in a few clicks!

#### Suggested Search Configurations

If you're new to building RAG systems, we strongly recommend selecting one of the suggested search configurations to start. These are high quality RAG setups that can give you state of the art performance. Simply select one of the following templates from "Docs & Search" > "Manage Search Tools" > "Add Search Tool":

* **Best Quality**: The best quality search configuration we’ve found. Uses Gemini 2.5 Pro to extract documents to text.
* **Cost Optimized**: Still excellent, but lower cost. This configuration uses Gemini 2.5 Flash to extract documents to text.
* **Vector Only**: A configuration which only uses vector search for semantic similarity, without keyword search. Useful when you want to search only on semantic meaning without weighing the keywords in the query.
* **OpenAI Based**: We suggest using a Gemini-powered config above if possible — they support more document types and have better document extraction quality. However, if you are required to use OpenAI APIs, try this configuration with GPT-4o. This config does not support transcribing audio and video.

We’re working on adding more document extractors and embedding providers to expand this list.

#### Search Tool Name & Description

When you’re creating a search tool, you’ll be asked to provide a tool name and description. These are important as the model will read them to decide if and when to use your search tool.

For example:

* **Poor:** search\_tool - “Search documents for knowledge”
* **Better:** doc\_search - “Search the knowledge base for product information”
* **Best:** knowledge\_base\_search - "Search Kiln's user-facing documents, guides, and walkthroughs."

In the first example, the model will have no idea what type of documents it has access to, or if/when to search them. The last example is much better; from it the model knows what the documents are, the audience they are written for, and can infer when searching them would be helpful to a task.

#### Custom Search Configurations

If you have experience with RAG systems, you can create a completely custom RAG. Simply select “Create Custom” on the “Add Search Tool (RAG)” screen.

{% hint style="warning" %}
**Advanced Users Only**: unless you have experience with AI embeddings and search, we suggest sticking to the suggested search tool configurations.
{% endhint %}

You can customize:

* Extractor: The model used to extract non-text documents (e.g. PDFs, videos) into text. Optionally customize the prompts passed to the extractor model for each type of file.
* Chunking Strategy: specify how large documents are split into smaller chunks for embedding, indexing, and retrieval
* Embeddings: specify the embedding model and embedding dimensions
* Search Index / Vector Store: select the search index strategy including vector index, full-text search, or hybrid mode.

Want to see more options here? Let us know on our [Discord](https://kiln.tech/discord)!

#### Processing Documents

After adding documents, Kiln must process them before they can be searched. You can monitor progress on the "Search Tools (RAG)" page. See [how it works](#how-it-works) below for more information about what each step is doing.

<figure><img src="/files/AHStF5tC082ShDmgpR3i" alt="" width="187"><figcaption><p>Waiting on Processing</p></figcaption></figure>

### Using Search Tools

Once you’ve [created a search tool](#building-a-search-tool) and [processing is complete](#processing-documents), you can run your search tools!

#### Using a Search Tool in a Task

To use a search tool in a task, simply select it from the “Tools” dropdown in the “Advanced” section of the Run page then run your task.

<figure><img src="/files/q3w6YlcICylWDiHKui8q" alt="" width="375"><figcaption><p>Selecting a search tool in "Run"</p></figcaption></figure>

The search tool will be provided to your task, and your model may invoke it. You can view the model’s tool calls and the search tool’s response in the “All Messages” section of the run page:

<figure><img src="/files/RCc2iPN0OJWvfZVBj8Gx" alt="" width="375"><figcaption><p>A trace including a tool call to a search tool</p></figcaption></figure>

{% hint style="info" %}
It’s the model’s choice if and when to invoke a tool. If the model isn’t invoking search when you feel it should, see the section below on [optimizing your RAG](#step-1-optimize-search-tool-name-description-and-task-prompt).
{% endhint %}

#### Testing Your Search Tool

If you just want to test your tool to see what it returns, you can do so from "Docs & Search" > "Search Tools" > search tool details. Enter any query to see what your search tool returns.

This mode is intended only for testing. It will render the raw chunks as would be returned to the AI task. You wouldn't normally expose these results directly to a user, and instead would have an AI task extract answers or summarize the content.

<figure><img src="/files/2zTdfAy5BYIa1PY1BQKL" alt="" width="375"><figcaption><p>A search tool test invocation</p></figcaption></figure>

### How it Works

Under the hood, there are 4 stages to Kiln's RAG/search pipeline:

1. Document Extraction: convert documents like PDFs, videos, and audio into text data that language models can read.
2. Chunking: break down large documents into smaller chunks
3. Embedding: generate semantic embeddings from your chunks
4. Search: index the embeddings and chunks in a vector database, then search it

### Optimizing your RAG

Kiln offers several options for improving your RAG. To do so, you can create multiple search tools, then compare their quality using either:

* Manually review search result quality in the [search tool test UI](#testing-your-search-tool)
* Write [evals](/docs/evals-and-specs/evaluations) to measure resulting task quality

{% hint style="success" %}
Kiln will minimize processing where possible. For example, if many search tools all share the same extraction config, it will reuse the prior extractions. This makes experimentation faster and reduces costs.
{% endhint %}

#### Step 1: Optimize Search Tool Name, Description and Task Prompt

Often we see issues where the search tool can easily retrieve the needed data, but the tool is never called. This is easy to identify: check the "All Messages" section of the run to see if the tool was invoked.

This is usually an easy fix with one of the following:

* Make the search tool name and description more descriptive: [example and guidance](#search-tool-name-and-description).
* Make the task's prompt explicitly define when search tools should be used, for example by adding "Always confirm answers with the knowledge\_base\_search tool."

#### Step 2: Improve Document Extraction

The first step of RAG is extracting your documents (PDFs, images, videos) into text which we can index, search, and provide to tasks after retrieval.

{% hint style="info" %}
If the data produced during extraction isn’t high quality, there’s nothing the rest of the pipeline can do to recover.
{% endhint %}

Walk through these steps to identify and improve document extraction:

1. **Inspect Extractions for Issues**: Read document extractions and compare with the original documents. You can do this by clicking on documents in the document library. Once you identify issues, you can fix them using the steps below. Example issues:
   * Including irrelevant data, like a header/footer content, transcribing menus/navigation content, or transcriptions of images which are embedded but not part of the core content (navigation, headers, even web ads).
   * Skipping important data, like insights from chart images
2. **Upgrade Extraction Model**: If you have issues, consider a higher quality extraction method. Often a better model will resolve extraction issues. We suggest trying Gemini 2.5 Pro via the Gemini API. While these APIs can be costly, you only need to extract documents once so it’s not a recurring cost.
3. **Customize Extraction Prompts \[Video Example Below]**: The default extraction prompts in Kiln are generalized prompts designed to work with any document. However, if you know your documents are a specific format, you can improve extraction by creating custom extraction prompts for your use case. You can do this when creating a new Search Tool, in the “Advanced section” of the extractor. See the demo video and the "Custom Extraction Prompt Examples" section below for details.
4. **Fully Custom Extraction**: If desired, you can always extract your documents separately using your own code, then add text (`.txt`) or markdown (`.md`) files to Kiln. This gives you complete control. Kiln won’t re-process files that are already in text/markdown formats.

{% embed url="<https://vimeo.com/1138970149>" %}

<details>

<summary>Custom Extraction Prompt Examples</summary>

Here is an example prompt if you know all documents are PDFs of blog posts:

{% code overflow="wrap" %}

```
Only transcribe the blog title, subtitle, byline and blog post content.

Ignore other elements of the page including headers, footers, and navigation elements.

Extract the blog data into markdown. Specify the post title as H1 (`#`) at the top, and all section titles should be smaller (H2, H3, etc).
```

{% endcode %}

Example for extracting invoices:

{% code overflow="wrap" %}

````
You are extracting invoice PDFs. Only extract payee, payer, date, status, invoice number, and invoice line items. Disregard headers, footers, addresses, and decorative elements.

Extract the invoice data into the following format:
```
Payee Name: X
Invoice ID: X
...
```
````

{% endcode %}

{% hint style="info" %}
Videos, Documents and Audio have separate prompts, so you can customize each to the use case as needed.
{% endhint %}

</details>

#### Step 3: Tune Chunking Size and Top-K

When your task calls your search tool, it will fetch a certain number of document chunks. Chunks are created by splitting long documents into smaller pieces. This is important for 2 reasons:

* You don’t want to feed too much information into the task, as it will flood the context, produce poorer results, and cost more.
* Splitting into chunks improves search relevance. A 50 page document might contain information on many topics. Searching for smaller chunks reduces the topic per segment, which helps your search tool find the most relevant portions of the document.

{% hint style="info" %}
The chunk size is defined when creating a search tool, in the chunking method options. You can also define how much overlap there is between chunks. The default is 512 words/tokens per chunk with 64 words/tokens overlap.

The number of results returned is called top-k, and is defined when creating a search tool, in the search index options. The default is to return 10 chunks.
{% endhint %}

Tuning these two variables for your use case can help produce better search results.

**Option 1: Increase Chunk Size and Reduce Top-K** Sometimes you know there’s exactly one document which will contain the answer; for example for the question “What is the total on invoice INV-123456?”. Returning 10 invoices won’t help this query, and will splitting the one invoice across 5 chunks could harm its performance. In this case, a larger chunk size and a small top-K would be a great configuration. You’ll still end up returning a reasonable amount of data, as you’ve lowered top-K.

**Option 2: Lower Chunk Size and Increase Top-K** Sometimes you know the model will need many of chunks to get a good answer; for example “Which protein structures were rated as ‘promising’ in experiments from June to July 2025?” might need to return hundreds of data chunks. In this case a small chunk size and higher top-K could work well.

{% hint style="info" %}
It's almost never a good idea to set Top-K to 1. There's always a chance that an answer is split across 2 or more chunks, so returning multiple chunks is always a good idea.
{% endhint %}

#### Step 4: Tune Search Index Options

Kiln has powerful search options, backed by [LanceDB](https://lancedb.com/):

* **Vector Search**: Searches for chunks based on an embedding/vector representation of their semantic meaning. This lets you find results that *mean* the same thing, even if the query uses completely different wording.
* **Full-text search (aka keyword search or BM25)**: Searches for literal words/terms. It scores chunks based on how often your keywords appear (term frequency) and how rare they are across the entire dataset (inverse document frequency).
* **Hybrid search**: Combines both vector and full-text search, giving you relevance by meaning *and* by exact keyword match.

{% hint style="info" %}
Vector and Hybrid search require calculating an embedding of the search query. This requires calling an embedding model which can take time, and cost money if using a paid API.
{% endhint %}

You can read more about search indexing and retrieval in the [LanceDB docs](https://www.lancedb.com/docs/search/) or on the [LanceDB Blog](https://blog.lancedb.com/hybrid-search-combining-bm25-and-semantic-search-for-better-results-with-lan-1358038fe7e6/).

We typically recommend **hybrid search**, but your use case might benefit from other options:

* **Full-text only**: best for cases where you want *exact term matching* (e.g. legal text search, log file search), or extremely fast performance.
* **Vector-only**: best for cases where *meaning matters more than exact words* (e.g. semantic question answering, summarization datasets).
* **Hybrid**: best for cases where you want both — i.e. match the meaning but still boost exact matches.

#### Step 5: Explore Embedding Models

As a last step, you can try different embedding models: the models which generate a embedding/vector-representation from a chunk.

Generally, we suggest exhausting the options above before tuning here.

### Deploying your RAG

Once you've optimized your RAG in Kiln, you're ready to deploy it!

You have several deployment options to choose from, depending on your use case:

#### **Kiln UI: For Personal Use**

You can continue to use the Search Tool inside Kiln, using the "Run" UI. This option is great for a single user or small teams. See our [collaboration docs](/docs/collaboration) for how to share a search tool with your team.

#### **MCP: For Local LLM Clients**

If you prefer another LLM frontend like LMStudio or Jan, you can run your Kiln Search Tool as an MCP server, then connect to it from your client of choice. See our [MCP server documentation for instructions](https://github.com/Kiln-AI/Kiln/tree/main/libs/server/kiln_server/mcp#readme) on running an MCP server exposing Kiln Search Tools.

#### **LlamaIndex: For Production Applications**

You can load your Kiln RAG dataset into a production-ready [LlamaIndex](https://www.llamaindex.ai/) stack. See our [Python library docs](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#taking-kiln-rag-to-production) for how to load a Kiln Search Tool into any LlamaIndex vector store.

<a href="https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#taking-kiln-rag-to-production" class="button primary">Deploy a Kiln Search Tool</a>

***We recommend*** [***LanceDB Cloud***](https://lancedb.com/) ***for production hosting***. The Kiln app and library use LanceDB; using LanceDB in prod ensures your production stack matches the RAG you optimized in Kiln perfectly. It is a performant vector database for any scale.

{% hint style="success" %}
Note: Loading a production index will not need to repeat extraction, chunking, and embeddings. Those steps are already completed in Kiln, and their results are saved in your Kiln dataset.
{% endhint %}


# Agents

Building Agentic AI systems has never been easier

<figure><img src="/files/xmTT57nE4cLypCgKGrOT" alt=""><figcaption></figcaption></figure>

Ah "agents", the most overloaded term in AI! The word "Agents" means different things to different people, but the good news is Kiln supports all the patterns typically associated with "agentic systems". Let's break down the common agent features, and how to build them in Kiln!

* [Tool Use](#tool-use)
* [Multi-Actor Interaction (aka subtasks)](#multi-actor-interaction-aka-subtasks)
* [Goal Directed, Autonomous Looping & Reasoning](#goal-directed-autonomy-and-reasoning)
* [State & Memory](#state-and-memory)

### Tool Use

Almost every definition of agents includes tool use. This is some way for the agent to interact with the outside world, such as calling a database, calling an API, or even sending an email.

Kiln has full support for adding tools to your Kiln tasks. Tools can be added to Kiln via [MCP servers](/docs/tools-and-mcp#connecting-tools), [Kiln Search Tools (RAG)](/docs/documents-and-search-rag), or [Code Tools](/docs/tools-and-mcp/code-tools) you write yourself in Python. See the [Tools & MCP](/docs/tools-and-mcp) docs for details.

### Multi-Actor Interaction (aka subtasks)

Agentic systems often involve an AI agent delegating/coordinating other agents. There are many variants of this pattern with names like orchestrators, supervisors, directors, executors, and experts/assistant.

Kiln has full support for this pattern via Kiln Task as Tools:

#### Kiln Tasks as Tools

In Kiln you can make any Kiln Task into a tool which other tasks can then call. This gives you the flexibility to create multi-actor patterns by organizing a hierarchy of tasks:

<figure><img src="/files/MPUw0MhRVi7zEsDNYydl" alt="" width="375"><figcaption><p>A Kiln task, able to call other Kiln tasks as tools</p></figcaption></figure>

To create a Kiln Task as Tool follow these steps:

**Step 1: Create a Kiln Task**

First create a new task. Select "New Task" in the project menu in the sidebar, and fill out the form with a name and prompt.

<figure><img src="/files/JHQJkRGmLKl2ri4EhTeP" alt="" width="275"><figcaption></figcaption></figure>

**Step 2: Add Tools**

If you haven't already and this task requires tools, add tools to your project following the instructions in our [Tools & MCP docs](/docs/tools-and-mcp).

**Step 2: Create a Run Configuration**

Each Kiln task is run with a "run configuration" which specifies which model is used, which prompt is used, which tools are allowed, and more. Since you want your tool calls to be completely consistent, you must first create and save a run configuration.

The easiest way to create a run configuration is from the "Run" tab in Kiln. You can tweak every parameter, run the task to ensure it works, then save it when you're happy.

<figure><img src="/files/pPt5yBKSLmqtMJaBDP7I" alt="" width="375"><figcaption><p>Creating a run configuration in the "Run" tab</p></figcaption></figure>

**Step 3: Create a Kiln Task as Tool**

To turn your task and run config into a tool other tasks can call, open "Tools" > "Add Tools" > "Kiln Task as Tool".

When creating you'll be asked to provide a number of options

* Kiln task: which task this tool will call
* Run configuration: which model to use, and which tools to allow this subtask to call
* Tool name: the name of the tool, which the calling task will see
* Tool description: the description of the tool, which the calling task will see

{% hint style="warning" %}
You’ll be asked to provide a tool name and description. These are very important as the model will read them to decide if and when to use this subtask / agent.

For example:

* **Poor:** search - “Search for information” (what will it search?)
* **Better:** web\_search - “Search the web for information” (does it return webpages? summaries? search listings?)
* **Best:** web\_researcher - "This tool can search and scrape the web for information. Ask it a question and it will research using the web, and reply with a summarized answer."
  {% endhint %}

<figure><img src="/files/SikdxP5EKuogAf0YVWtx" alt="" width="375"><figcaption><p>Creating a Kiln Task as Tool</p></figcaption></figure>

**Step 4: Use Your Task as Tool**

To use your new subtask, simply select it from the "Tools & Search" dropdown on the Run tab.

<figure><img src="/files/DhhWxab5dXH6pZWhMFvP" alt="" width="375"><figcaption><p>Adding a Kiln task as tool subtask</p></figcaption></figure>

If you always want this task to have access to this tool, use the "Save current options" and "Set as task default" links to create a default run configuration including this task.

**Step 5: Viewing Tool Calls**

When you run a task that has access to subtasks, you're able to view the tool/task invocations in the "All Messages" list. Click the "Messages" link to see the subtasks message list.

<figure><img src="/files/nQ4siEwvet56AlfkreWm" alt="" width="375"><figcaption></figcaption></figure>

#### Context Management

A common issue with agentic systems is that the context window (chat history) gets very long. This can cause several major issues:

* **Filling the context window**: if the context window grows larger than the model supports, it will error and fail
* **Degraded quality**: a large context window can degrade the quality of newly generated content, especially if some of the messages are no longer relevant or inaccurate.
* **Increased costs**: each new message has to process all of the tokens in the chat history. Even a short message at the end of a long chain can be expensive. \[note]

{% hint style="success" %}
**Using subtasks is the easiest way to solve context management issues!**
{% endhint %}

Let's look at a visual example of an agent which requires many web-searches to build a final answer:

* **Without Sub-Agents**: the context becomes very large very quickly. Entire webpages are loaded into context and stay there. Data we don't need, like web pages that didn't yield useful information, keep taking context room forever. Later calls need to process a large number of tokens to generate the next new token (increasing cost).
* **With Sub-Agents**: each web-research task is performed in its own sub-agent. The results of the webpages are summarized and only the important details are returned to the main task. The context of the main task stays focused and small. Each subtask is also smaller and more focused. Irrelevant data is dropped when the subtask ends. Costs are lower by approximately a factor of 4.

<figure><img src="/files/B52CEjQ6Vk0Q6a1l0RDK" alt=""><figcaption></figcaption></figure>

### Goal Directed, Autonomy, & Reasoning

Agents typically aim to achieve a specific goal, and are given the ability to run over time to achieve that goal.

* **Goal-directed**: acts to achieve objectives rather than just replying
  * **In Kiln:** you define the agent's goal when you create the task prompt, and when you pass the user input
* **Planning & reasoning:** agents can break problems into steps, decide what to do next, and reason between tool calls
  * **In Kiln:** All tasks plan internally about which tools to call. You can tell your tasks to perform additional reasoning/planning by either 1) selecting a reasoning model, 2) selecting a "Chain of Thought" prompt which will explicitly ask the model to reason during the run loop.
* **Autonomy / Looping**: the agent can choose how to achieve the goal, calling any tool in any order. They can loop, mixing thinking/reasoning and tool calls until their task is complete.
  * **In Kiln:** Every Kiln task can loop, making many tool calls until deciding to return a final result. The task has autonomy to decide which tools to call, in which order.

{% hint style="success" %}
**Advanced Users: 'ReAct' in Kiln**

A common agent pattern is named 'ReAct'. It was defined in a paper by [Yao et al](https://arxiv.org/abs/2210.03629).

While we don't use the name ReAct in the Kiln app, a Kiln task with a chain-of-thought prompt has the same reasoning+acting steps as described in the ReAct paper.
{% endhint %}

### State & Memory

Agents maintain state over steps, letting them make progress towards their goals.

Currently Kiln maintains memory simply through the message history. See the [context management](#context-management) section above for how to use Kiln subtasks to compress memory, summarizing many subtask messages into a shorter summaries in the main-task history.

We're exploring additional memory management options in Kiln. If you have requests, please let us know on the [Discord](https://kiln.tech/discord)!


# Skills

Load more instructions, depending on the goal (Progressive Prompt Enhancement)

Skills are reusable instructions that your agents can load on-demand. Unlike prompts which are always included in your agent's context, skills are only retrieved when the agent decides they're needed. This makes skills perfect for specialized knowledge, detailed procedures, or domain expertise that you don't want cluttering every single run.

### Creating Skills

Skills are created at the project level, and can be used by any task within that project.

To create a new skill, navigate to the "Skills" page in your project sidebar and click "New Skill".

You'll be prompted to provide:

* **Name**: An identifier for the skill. Only allows lower case letters, numbers and hyphens.
* **Description**: A clear explanation of what the skill does and when to use it
* **Instructions**: The markdown-formatted instructions that will be loaded when the agent calls this skill.

{% hint style="warning" %}
**Write Clear Descriptions**

The skill description is what the model sees when deciding whether to use a skill. Make it specific and actionable:

* **Poor:** "SQL help"
* **Better:** "Write and optimize SQL queries for PostgreSQL"
* **Best:** "Generates SQL queries for PostgreSQL with proper joins, filtering, and performance optimization. Use when user requests data retrieval, filtering, or analysis from a database."
  {% endhint %}

### Using Skills

Once you've created skills, you can make them available to your agents through run configurations.

#### Adding Skills to Run Configurations

On the "Run" screen, under the "Skills" section in Advanced, you'll see all available skills in your project. Select the skills you want to make available to the agent for this run.

#### How Agents Use Skills

When you make skills available to an agent, here's what happens:

1. **Discovery**: The agent sees all available skill names and descriptions at the end of its system prompt.
2. **Decision**: The agent decides which skills (if any) are relevant to the current task.
3. **Loading**: When the agent needs a skill, it calls the `skill()` tool with the skill's name.
4. **Execution**: The skill's full instructions are loaded into the conversation, and the agent proceeds with the task.

This lazy-loading approach means your agent's context stays clean and focused. Specialized instructions are only retrieved when actually needed.

{% hint style="info" %}
**Model Behavior**

The model may or may not choose to use the skills you provide. If you find the model is not using a skill when you feel it should, try:

1. Making the skill description more specific and actionable
2. Updating your prompt to explicitly mention when certain skills should be used
3. Ensuring the skill name clearly indicates its purpose
4. Using a more modern model: models that came out before Agent Skills were released have not seen skills in their training data, and are less likely to leverage them.
   {% endhint %}

#### Viewing Skill Usage

When viewing a run trace, you can see which skills the agent loaded in the "All Messages" section. Skill calls appear as tool invocations, showing when and how the agent used each skill.

### Reference Files

Skills can include reference files—additional markdown documents that provide supporting information, which are also only retrieved when the agent decides they're needed. References are useful for:

* Detailed technical specifications
* Code examples or templates
* Domain knowledge or terminology
* Style guides or formatting rules

#### Adding References

Currently, references are managed through the file system. To add references to a skill:

1. Navigate to your project's `skills` directory
2. Find the skill's folder (named after the skill ID)
3. Create a `references` folder if it doesn't exist
4. Add markdown files (`.md`) to the `references` folder
5. Reference the new file from your main SKILL.md file (see details below).

Your skill structure should look like:

```
skills/
  <skill_id>/
    skill.kiln           # Metadata (name, description, etc.)
    SKILL.md             # Main skill instructions
    references/          # Additional reference files
      API_RESPONSE_FORMAT.md
      STYLE_GUIDE.md
      ...
```

#### Referencing from SKILL.md

Within your skill instructions (in `SKILL.md`), you should reference files using markdown link syntax. This is how the agent can discover reference files. Example:

{% code overflow="wrap" %}

```markdown
## Instructions

Follow our company's coding standards. See our [style guide](references/STYLE_GUIDE.md) for details.

When generating API responses, use the format specified in this [API format guide](references/API_RESPONSE_FORMAT.md).
```

{% endcode %}

### Skills vs Tools vs Subtasks

Kiln offers three ways to extend your agent's capabilities. Here's how to choose:

<table data-full-width="true"><thead><tr><th>Feature</th><th>Skills</th><th>Tools</th><th>Subtasks</th></tr></thead><tbody><tr><td><strong>Primary Use</strong></td><td>Instructions &#x26; knowledge</td><td>External actions</td><td>Complex workflows</td></tr><tr><td><strong>Context Impact</strong></td><td>Details loaded only when needed. Only description injected into prompt.</td><td>Tool definitions often much more verbose than skills descriptions</td><td>Isolated per subtask</td></tr><tr><td><strong>Progressive Disclosure</strong></td><td>Yes - can link to reference files, forming a nested hierarchy of knowledge</td><td>No</td><td>No</td></tr><tr><td><strong>Best For</strong></td><td>Guidelines, procedures, rules</td><td>APIs, databases</td><td>Multi-agent patterns</td></tr><tr><td><strong>Setup Complexity</strong></td><td>Low</td><td>Medium</td><td>High</td></tr></tbody></table>

### Skills vs Search Tools (RAG)

Skills can also be compared to [RAG](/docs/documents-and-search-rag). Both have pros and cons for adding knowledge to a system:

<table data-full-width="true"><thead><tr><th>Feature</th><th>Skills</th><th>RAG</th></tr></thead><tbody><tr><td><strong>Max Documents</strong></td><td>Typically &#x3C;50</td><td>Unlimited</td></tr><tr><td><strong>Context Impact</strong></td><td>Low: ~100 tokens per document for name and description.</td><td>Very Low: 1 tool definition for all documents.</td></tr><tr><td><strong>Chunking</strong></td><td>Manually create skills and reference files.</td><td>Automatic chunking of larger docs.</td></tr><tr><td><strong>Control</strong></td><td>High: define the exact skill and reference file content. Descriptions can prompt the agent exactly when to load each skill.</td><td>Low: documents are automatically chunked (split) and indexed. The agent needs to guess search terms to find relevant information.</td></tr><tr><td><strong>Deployment Complexity</strong></td><td>Low: no additional effort if your toolchain supports it (like <a href="/pages/2MUphEoXwKJgVQbP1CqJ">Kiln SDK</a>).</td><td>High: requires vector database hosting, and jobs to process and index documents.</td></tr><tr><td><strong>Progressive Disclosure</strong></td><td>Yes - can link to reference files, forming a nested hierarchy of knowledge</td><td>Yes - can muli-hop.</td></tr></tbody></table>

### Best Practices

#### Write Specific, Actionable Descriptions

Your skill descriptions are the primary signal the model uses to decide when to use a skill. Make them clear and specific:

* ✅ "Validates JSON schemas against the OpenAPI 3.0 specification and reports errors"
* ❌ "JSON validator"

#### Keep Skills Focused

Each skill should have a single, clear purpose. If you find a skill becoming too broad, consider splitting it into multiple focused skills, or splitting the skill into [reference files](#reference-files).

#### Test Your Skills

After creating a skill, test it by:

1. Creating a run configuration with the skill enabled
2. Running your task with inputs that should trigger the skill
3. Checking the trace to verify the skill was loaded and used correctly

### Managing Skills

#### Editing / Cloning Skills

We don't allow editing skills, as historical eval scores and dataset runs become would become misleading if edited, with no versioning to tell you what changed.

Instead, Kiln lets you clone a skill, giving you a fresh copy to modify while keeping the original intact so existing results stay valid.

#### Archiving Skills

If you have skills you no longer use but want to keep for reference, you can archive them. Archived skills won't appear in the skills dropdown and won't be available to agents. Unarchiving restores them to active use.


# Synthetic Data Generation

Generate synthetic data for fine-tuning or evaluation

Anyone can create thousands of synthetic data samples in just a few minutes using our interactive UI.

<figure><img src="/files/wKlx61fGu9bta0qP6a0l" alt=""><figcaption></figcaption></figure>

## Use Cases

Synthetic data is helpful for many reasons:

* **Evals:** Generate data for custom evals of your task performance
* **Fine-tuning:** Generate fine-tuning datasets
* **Built-in Quality Templates:** Use our built-in data-gen templates like 'Jailbreak' or 'Bias' to check your system for common issues (curated evals)
* **Addressing Bugs / Issues:** generate targeted data to reproduce a bug/issue, which can be used for training a fix, evaluating a fix, and backtesting
* **Prompting:** Generate examples to be used for few-shot or multi-shot prompting

## Get Started

There are 2 portions of synthetic data in Kiln:

* [**Synthetic Data Guide**](/docs/synthetic-data-generation/synthetic-data-guides) **(optional):** Kiln learns what great data looks like, to generate even better data for your task.
* [**Generating Synthetic Data**](/docs/synthetic-data-generation/generating-synthetic-data): build synthetic datasets for evals or fine-tuning


# Synthetic Data Guides

Improve your synthetic data generation

A Data Guide is a per-task prompt that tells Kiln what realistic **inputs** to your task look like: their structure, style, terminology, and value ranges. Without one, the data generation model has to guess what your domain looks like from the system prompt alone. With one, generated topics and inputs are shaped to match your actual data.

<figure><img src="/files/G60rFJ5bkaFFt7YTnmAo" alt=""><figcaption></figcaption></figure>

### When to Use Data Guides

Most tasks benefit from a Data Guide, especially when:

* Your inputs have a specific structure (forms, JSON, transcripts, tickets, code) the model would otherwise have to invent
* Your domain uses terminology, value ranges or constraints a generic model wouldn't know
* Default synthetic data doesn't look like production data
* You already have real examples (documents, past runs, or a spreadsheet of inputs) and want generation grounded in them

#### Setting Up A Data Guide

The first time you open Synthetic Data Generation for a task, Kiln offers to set up a Data Guide. Click **Set Up Data Guide** and Kiln will ask how you'd like to build it: **Manually**, or with **Kiln Pro**.

<table><thead><tr><th valign="middle"></th><th valign="middle">Manual</th><th valign="middle">Kiln Pro</th><th data-hidden></th></tr></thead><tbody><tr><td valign="middle"><strong>AI Guided Authoring</strong></td><td valign="middle">Manual</td><td valign="middle">Automatic</td><td></td></tr><tr><td valign="middle"><strong>Style &#x26; Constraints Discovery</strong></td><td valign="middle">Manual</td><td valign="middle">Automatic</td><td></td></tr><tr><td valign="middle"><strong>Learn From Documents</strong></td><td valign="middle">—</td><td valign="middle">✅</td><td></td></tr><tr><td valign="middle"><strong>Approx. Effort</strong></td><td valign="middle">~15 mins</td><td valign="middle">~5 mins</td><td></td></tr><tr><td valign="middle"><strong>Kiln Account</strong></td><td valign="middle">Optional</td><td valign="middle">Required</td><td></td></tr></tbody></table>

Both paths end in the same place: a saved guide you've reviewed and refined. The difference is how the guide gets written.

#### Kiln Pro Data Guide Generation

With Kiln Pro, you provide Kiln a set of real example inputs and it analyzes them into a complete data guide for you. It's the fastest path, and the only one that can learn directly from documents.

{% hint style="success" %}
Kiln Pro Data Guides learns from many examples to generate the best possible synthetic data. It learns what's consistent between documents, what varies, relationships, formatting, style and more. It's the best way to create high quality syntehtic data.
{% endhint %}

1. **Connect Kiln Pro** (if you haven't already).
2. **Add example inputs.** Click **Add Inputs** and pick a source: 1) upload documents from your computer or reuse them from your [Document Library](/docs/documents-and-search-rag), 2) pull inputs from real past runs in your [dataset](/docs/organizing-datasets), 3) bulk-import from a CSV (plaintext tasks only), or 4) write one by hand (structured tasks only). You can mix sources and add more later.

{% hint style="info" %}
Real data works best. A Data Guide is only as realistic as the examples it's built from, so prefer real inputs over synthetic ones.
{% endhint %}

#### Manual Data Guide Generation

Manual guides need no Kiln account. You start with a few examples of existing data, then refine the guide after seeing how it performs.

1. **Add example inputs:** provide at least one real input, either typed in or picked from your existing task runs. Your examples are treated as reference data for future data generation.
2. **Generate a preview:** Kiln produces a set of synthetic inputs from your guide so far, and you refine from there.

#### Review and Refine

Both Pro and Manual flows include a refinement phase. Kiln will generate synthetic examples from your guide, which you can review and give feedback on. If you rate any **Needs Work**, Kiln will work with you on refining and improving your data guide. Once all the data is consistently being rated **Realistic** your data guide is complete and ready to be saved.

### Using A Saved Data Guide

Once saved, Kiln applies your guide automatically when [generating inputs for your task](/docs/synthetic-data-generation/generating-synthetic-data). A **Use Data Guide** toggle lets you turn it off for individual runs.

If you're also using synthetic data guidance (like from an eval), that guidance takes priority over the Data Guide where the two conflict.

To view, edit, or delete a saved guide, use the **Data Guide** button on the Synthetic Data Generation page.


# Generating Synthetic Data

Improve evals or fine-tuning with synthetic data

<figure><img src="/files/wKlx61fGu9bta0qP6a0l" alt=""><figcaption></figcaption></figure>

## How It Works

Kiln doesn't require you to write complex custom synthetic data gen prompts. Since you've already defined a goal when setting up your task, Kiln can do this for you. It will infer the type of data needed from the system prompt, adapt it to your data-gen goal, and create synthetic data gen prompts without any manual prompting.

Kiln offers two ways to build the dataset itself:

* [**Kiln Pro Batch Planning**](#kiln-pro-plan-the-batch)**:** describe the dataset you want and Kiln plans the whole dataset for you, writing one tailored prompt per sample to cover your task's use cases and edge cases.
* [**Manual Generation**](#manual-build-a-topic-tree)**:** plan the coverage yourself by building a tree of topics for breadth.

Both flows start the same way: choose a goal, optionally set up a Data Guide, then pick how to build the dataset.

## Choose A Goal

First select a goal for your dataset generation: **Evals** or **Fine-Tuning**. This is an important step as you need different data for different goals:

* **Fine-Tuning**: generate high quality outputs across a broad range of possible inputs, to help your model learn how to respond to a range of requests. This can include generating inputs that commonly produce issues, and outputs that avoid that issue.
* **Evals**: Intentionally generate a mix of good and bad inputs and outputs. We'll use the bad outputs to ensure the judge model can properly assess failures, and we'll use the bad inputs to ensure your task no longer has the issue.

Selecting the goal will set up two properties:

* **Template:** targets the data generation to the use case you selected above. Use our built in templates for fine-tuning or evals, or generate your own custom guidance.
* **Tag Assignments:** which dataset tags will be assigned to generated data. This could be a single tag like `fine_tuning_data` or a randomly assigned split like `eval_data: 80%, golden_data: 20%`. These will be pre-filled based on your selected goal.

## Choose A Data Gen Model

{% hint style="info" %}
**TL;DR:** Choose a high quality model like the latest GPT or Claude model for synthetic data gen. Synthetic data gen is complex, and benefits from larger models.
{% endhint %}

We highly recommend choosing a large capable model for data gen. While your task may work on smaller models, data gen is more complex. It requires reasoning about a range of possible inputs, probing edge cases, and more. It benefits from a large model with a long context.

If generating content to evaluate how your model responds to inappropriate requests (bias, jailbreaking, maliciousness, etc.), choose an uncensored model like Grok or Dolphin. Censored models like GPT will refuse to generate some types of sensitive content.

## Building The Dataset

With your goal and Data Guide in place, Kiln asks how you want to build the dataset: **Manually**, or with **Kiln Pro**.

<table><thead><tr><th valign="middle"></th><th valign="middle">Manual</th><th valign="middle">Kiln Pro</th><th data-hidden></th></tr></thead><tbody><tr><td valign="middle"><strong>Effort</strong></td><td valign="middle">~15 min</td><td valign="middle">~5 min</td><td></td></tr><tr><td valign="middle"><strong>Use Case Coverage</strong></td><td valign="middle">Manual</td><td valign="middle">AI Planned</td><td></td></tr><tr><td valign="middle"><strong>Edge Case Coverage</strong></td><td valign="middle">Manual</td><td valign="middle">AI Planned</td><td></td></tr><tr><td valign="middle"><strong>Kiln Account</strong></td><td valign="middle">Optional</td><td valign="middle">Required</td><td></td></tr></tbody></table>

Choose **Kiln Pro** when you want a dataset that covers your task's use cases and edge cases without designing that coverage yourself. Choose **Manual** when you want to shape the dataset topic by topic.

Both flows generate synthetic data in the same four stages, and you curate at each one. Only the **first** stage differs between them:

1. **Plan or Topics:** Kiln Pro [**plans the whole batch**](#kiln-pro-plan-the-batch) up front; manual mode [**builds a topic tree**](#manual-build-a-topic-tree) for breadth.
2. [**Inputs**](#generate-inputs): generate synthetic model inputs (the user message).
3. [**Outputs**](#generate-outputs): run your task on the inputs to generate synthetic outputs.
4. [**Save Data**](#save-your-data): save your curated data into your dataset for use in evals and fine-tuning.

Whichever way you build the dataset, generation is interactive. Be critical of the generated data and use the UI to make great quality data: remove anything that doesn't match your goals, add guidance to steer the content, and iterate until you're happy with the results.

### Dataset Planning

A common issue with synthetic data generation is that if you ask a model to generate synthetic data 1000 times, you get 1000 very similar outputs. They are too uniform to be useful for fine tuning or evals.

The solution is to design your entire dataset to ensure breadth and coverage.

Kiln offers 2 ways to plan entire deadsets: [Kiln Pro Batch Planning](#kiln-pro-plan-the-batch) and [Manual Topic Trees](#manual-build-a-topic-tree).

#### Kiln Pro: Plan the Batch

Both modes produce a batch of data. The difference is the planning. Instead of building a topic tree, you describe the dataset you want and Kiln plans the whole batch up front, writing one tailored prompt per sample so every sample has a distinct purpose. The plan is what covers your use cases and edge cases and gives the batch its diversity (the job a topic tree does in manual mode), and it's yours to review and edit in a single scannable overview before anything is generated. It's a streamlined, single-step path from idea to dataset.

**1. Connect Kiln Pro** (if you haven't already).

**2. Describe the batch you want.** The **Generate Synthetic Data Batch** page has three controls:

* **Sample Count:** how many samples to plan.
* **Guidance:** free text describing the dataset you want, e.g. *"10% of the dataset should be in Spanish. 30% should hit edge case X.".* Kiln will prefill this from your goal. This single Guidance box is where you steer the whole batch; unlike manual mode, there's no separate per-stage guidance.
* **Use Data Guide:** include your task's [Data Guide](#set-up-a-data-guide), so both the plan and the generated inputs match the shape of your real data.

**3. Review the batch plan.** Click **Generate Batch** and Kiln drafts a **Batch Plan** showing what it intends to generate, before it generates anything:

* **Batch Overview:** a short summary of what the batch covers and how it's distributed.
* **All Dataset Items:** expand this to read every planned prompt, one per sample. Remove any you don't want from the row's "..." menu.

#### Manual: Build a Topic Tree

In manual synthetic data generation, you shape the dataset by building a topic tree: a hierarchy of topic nodes to generate data samples for. You can use the topics and sub-topics to control the distribution of data in your dataset.

{% hint style="info" %}
This video walks through the manual flow and predates Kiln Pro batch planning. The UI has changed since recording, but the steps are similar.
{% endhint %}

{% embed url="<https://vimeo.com/1088940292>" %}
Synthetic Data Generation Walkthrough
{% endembed %}

Kiln can use AI models to generate a topic tree for you from your task's prompt. It uses the prompt to ensure the topics are relevant to your goal. See the example above: the model knew it was building topics for newspaper headlines and generated appropriate topics. To generate topics, click **Add Topics**:

You can nest sub-topics under any topic, forming the tree. Adding layers allows you to quickly generate a significant amount of diverse data. Open any topic's "..." menu to expose an **Add Subtopics** button:

You can manually add topics instead of using synthetic topic generation. Select the "or manually add topics" option at the bottom of the "Generate Topics" dialog.

Topics are strongly recommended, but are optional. You can skip topics and add model inputs without topics by continuing to the **Generate Inputs** step.

### Generate Inputs

Model inputs are the data passed into your task. When normally running your task, these would likely come from a human. However, in synthetic data generation we use AI models to generate them.

In manual mode, click **Generate Inputs** to produce inputs in your data table (under each topic, if you're using them). In Kiln Pro, click **Generate Batch** on your reviewed plan and Kiln generates every sample's input in parallel.

Review the quality of inputs and ensure you're happy with them before proceeding. You can remove individual inputs for manual curation or reset the session to change the guidance and generate again to get better quality data.

### Generate Outputs

Once you have generated all of the inputs you want, click **Generate Outputs** to run your task on each input:

Generating will result in an output for each input:

Review the quality of outputs and ensure you're happy with them before proceeding. You can remove individual outputs for manual curation or reset the session to change the guidance and generate again to get better quality data.

### Save Your Data

Use the Kiln synthetic data UI to review your data. Once you're happy with the data, click **Save All** to save it into your dataset for use in evals and fine-tuning.

The data will automatically be tagged with appropriate tags, based on the goal you selected ([see details](#tagging)):

Once saved, you can view all of your saved data in the Dataset tab.

### Templates and Custom Guidance

In manual mode, you steer each stage with its own **Guidance**: separate instructions for topic, input, and output generation. Kiln starts each from a template chosen from your goal (or the eval you came from), so you don't have to write a data-gen prompt from scratch. Switch templates or write custom guidance before running any stage.

The guidance dropdown offers a few families:

* **Built-in templates:** ready-made guidance for common goals like **Fine Tuning**, **Toxicity**, **Bias**, **Maliciousness**, **Jailbreak**, and **Factual Correctness**. Use these to probe your system for common issues (curated evals).
* **Eval template:** when you start data gen from a specific [eval](/docs/evals-and-specs/evaluations), Kiln auto-selects the template that matches that eval's type, so the data targets exactly what the eval measures. Depending on the eval you'll see one of **Requirements Eval**, **Desired Behaviour Eval**, **Appropriate Tool Use**, or **Issue Eval** (legacy evals). Only the one matching your eval is shown.
* **Custom:** write your own guidance from scratch.

These templates are a starting point. Edit one before running a stage, or write your own guidance from scratch, to get exactly the data you want.

Some examples of custom guidance:

* Generate content for global topics, not only US-centric
* Generate examples in Spanish
* The model is having trouble classifying sentiment of sarcastic messages. Generate sarcastic messages.

{% hint style="info" %}
Often custom guidance is used for producing adversarial content: poor quality or inappropriate content. This is done to ensure an [evaluation](/docs/evals-and-specs/evaluations) can detect and fail this sort of content.

However, LLMs will often do their best to avoid producing poor or inappropriate content, even when asked for it. If you find that's the case, use an uncensored and unaligned model like Dolphin or Grok. These models will follow instructions more closely, and do not attempt to censor their content.
{% endhint %}

{% hint style="info" %}
Kiln Pro works differently: it plans the whole batch from a single prefilled **Guidance** box, with no template dropdown or per-stage guidance. See [Kiln Pro: Plan the Batch](#kiln-pro-plan-the-batch).
{% endhint %}

## Structured Data (JSON, Tool Calling)

If your task requires structured input and/or output, your synthetic data generation will automatically follow the schemas you defined. All values are validated against the schemas you define, and nothing will be saved into your dataset if they don't comply.

## Tagging

All synthetic data will be assigned a series of [tags](/docs/organizing-datasets#using-tags-to-organize-your-dataset) in the dataset:

* The tag `synthetic` (manual and imported runs have their own tags)
* A unique tag to identify the data session (e.g. `synthetic_session_12345`)
* Custom tags. These are set up automatically when you select a goal, but you can edit them before generating data:


# Input Templates & Feature Engineering

Filter and format model inputs

Feature engineering is the process of controlling which data (features) you expose to a model. When you expose only what's needed — without irrelevant data or duplication — models often perform better. With LLMs, even the order and formatting of your data matters.

### Input Templates

Kiln makes it easy to experiment with feature engineering using input templates. They let you transform your raw input data into a new, cleaner format before it's sent to the model.

<figure><img src="/files/uFtLkdAUCD672UjG8avj" alt=""><figcaption></figcaption></figure>

#### Jinja Templates

Kiln lets you write [Jinja2 templates](https://github.com/pallets/jinja) to transform your input into other formats. Jinja has a range of powerful tools like filters, loops, conditionals, field access, and JSON helpers. AI agents like Claude Code and ChatGPT are excellent at writing these templates if you're new to Jinja.

For example, a template that pulls just two fields out of a larger structured input:

```jinja
Project Title: {{ input.title }}

Summary: {{ input.summary }}
```

To create an input transformer, go to the **Run** tab in Kiln, expand **Advanced**, and select **Create Template** under **Input Transformer**. You'll be prompted for a Jinja template. You can then run your task as usual.

You can optionally save this run config (including your template) for use in [evals](/docs/evals-and-specs), which will give you metrics on how much it helps (or hurts) task performance.

#### Template Inputs

Templates have access to exactly one variable: `input`, which holds your entire task input. Insert the whole thing with `{{ input }}`, or reach into it with attribute and index syntax. The same `input` reference works for every task type — there's no special unpacking to learn.

What `input` contains depends on your task's input type:

| Task input type     | `input` is… | Example template                                     |
| ------------------- | ----------- | ---------------------------------------------------- |
| Structured (schema) | The object  | `{{ input.question }} — {{ input.context }}`         |
| List / array        | The list    | `{{ input[0] }}`, `{{ input \| length }}`, or a loop |
| Plain text          | The string  | `{{ input }}`                                        |

If your task takes plain text and that text happens to be valid JSON, Kiln parses it automatically so you can use field access — an input of `{"name": "Alice"}` lets you write `{{ input.name }}`. If the text isn't valid JSON, `{{ input }}` echoes it verbatim.

**Technical notes**

* Input templates are stored as part of a run configuration. Your dataset doesn't need to change at all — all past dataset items remain valid, while still transforming what the LLM sees. Your callers still pass the exact same data as before, and your original input is always preserved on the run record; only the message the model sees is transformed.
* Built-in Jinja2 filters are available (`length`, `join`, `default`, `tojson`, `upper`, etc.), but custom filters are not. Execution is sandboxed — no inline Python code.

### Feature Engineering Techniques

Here's a (very) brief introduction to the basics of feature engineering.

#### Filtering

Filtering out unnecessary data is a great way to improve AI performance. Look for things you can filter:

* ID fields that have no meaning to the model (UUIDs, internal keys)
* Unnecessary formatting
* Irrelevant fields or data the model doesn't need for the task at hand

#### Formatting

* Convert from token-heavy formats like JSON down to just the data you need (see the example above).
* Convert XML/JSON data into data with better variable names (`ref_res` → `refund_reason`).
* Convert internal enums/codes into data the model can understand (`cccb` → `credit_card_chargeback`).
* Limit the length of plaintext fields to prevent unbounded input size. For example, truncate a description to 300 characters:

  ```jinja
  Description: {{ input.description[:300] }}
  ```
* Control the order information is presented. High-level → detailed is typically best, but API data can also be key-sorted or left unsorted.
* We've even seen performance improve by pretty-printing JSON!

{% hint style="success" %}
The best way to know whether a transformation helps is to measure it. Save your run config and compare it against the original using [Kiln Evals](/docs/evals-and-specs).
{% endhint %}


# Issues

Kiln Issues are an AI-native issue tracker for AI teams

<figure><img src="/files/NpU0tkPTRkcVgRjLsBL3" alt=""><figcaption></figcaption></figure>

Software teams have tools like Github Issues and JIRA for tracking issues. What should AI product teams use?

Kiln Issues are an AI-native issue tracker for AI teams. It doesn’t just collect bug reports, but keeps the structured data needed to reproduce, evaluate and fix AI system issues.

{% hint style="info" %}
**Under Development**: Kiln Issues are a new feature. Expect lots of improvements over the coming months!
{% endhint %}

### How are Kiln Issues Different than Github Issues, JIRA, etc?

Kiln issues are a way for teams to track known problems, just like Github Issues, JIRA, and countless other bug trackers.

Here's what makes Kiln Issues different from standard software bug trackers:

* **AI-Data Native**: Kiln Issues include the input/output data, model, provider, hyperparameters, and other data you need to diagnose and reproduce an issue
* **Evals Integration**: Issues include an eval that can measure if the issue is fixed or not
* **Synthetic Data Integration:** Kiln Issues can use data samples to generate synthetic data to reproduce the issue
* **Prompts:** Issues include prompts, not just freeform text discussions.
* **Longer lived:** Software bugs are usually closed once fixed. AI issues keep measuring over time, watching for regressions.

### Creating An Issue

To create an issue, select the Issue template when creating an Eval on the Eval tab. Include the following data:

* Prompt: A prompt describing the issue, which a judge model can use to determine if a model input/output pair exhibits the issue.
* Failure example (optional but recommended): an example of what failure looks like. Be sure to include this if the issue is subtle, subjective, or difficult to understand without examples.
* Passing example (optional)

<figure><img src="/files/g9cPrjOT80EI11SeNGBz" alt="" width="375"><figcaption><p>Create Issue UI</p></figcaption></figure>

### Evaluating An Issue

Issues are evals in the Kiln UI, but they go beyond a basic eval. Issue Evals have custom synthetic data gen templates, understanding that we want to generate inputs that reproduce a specific issue. It will help you easily create data samples from a single description (and optional examples).

Once you're done setting up your Issue Eval, you can use it to:

* Ensuring the issue is fixed: use the Issue Eval to confirm your fix (prompt change, model change, fine-tune) actually resolves the issue
* Ensuring the issue never regresses: keep running prior issue evals to ensure you don't accidentally regress the issue

Here's an example of the eval compare screen, ensuring several issues don't regress or are improved as they iterate/improve the AI system:

<figure><img src="/files/VYDOyZZVu6q0mZJkdlkg" alt="" width="375"><figcaption><p>Comparing several issues at one time</p></figcaption></figure>

### Philosophy: AI Product Evals work Best with Many Small Issue Evals <a href="#setup-team-evals" id="setup-team-evals"></a>

At Kiln we believe if creating an eval takes less than 10 minutes, your team will create them when they spot issues or fix bugs.

When evals become a habit instead of a chore, your AI system becomes dramatically more robust and your team moves faster.

Read our ~~manifesto~~ [guide on how to setup evals for your team](https://kiln.tech/blog/you_need_many_small_evals_for_ai_products#setup-team-evals). It covers:

* Many Small Evals Beat One Big Eval, Every Time
* The Benefits of Many Small Evals
* Evals vs Unit Testing
* 3 Steps to Set Up Your Team for Evals and Iteration


# Reasoning & Chain of Thought

Improve your model's quality with inference time scaling

<figure><img src="/files/oL2vN6b6H0k8lzrJxzPb" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Want to dive in and build a reasoning model? See our [Guide for Training A Reasoning Model](/docs/fine-tuning/guide-train-a-reasoning-model)
{% endhint %}

Kiln has powerful support for reasoning models and chain of thought. These techniques can generate higher quality results, while also reducing costs and improving performance.

<details>

<summary>What are reasoning models and chain of thought?</summary>

Reasoning models and chain of thought (COT) are methods that give models time to "think" before giving a final answer. Their "thinking" takes the form of discussing the request and possible answers in a stream of generated tokens. These additional tokens allow for more complex reasoning, step-by-step thinking, and have been shown to improve the quality of results.

These approaches are also known as "inference time scaling," where models improve from spending more compute power at inference time — as opposed to improving by spending more compute at training time.

While similar in some ways, the methods have some differences:

* **Chain of thought** is a method that's been around for a few years, and simply involves asking the model to think before giving an answer. This can be as simple as appending "Think step by step" to your prompt or adding detailed instructions for what the model should "think" about before giving its final answer.
* **Reasoning/thinking models** like Deepseek R1 or OpenAI's O3 are a newer form of inference time compute, where the model itself was trained to develop powerful reasoning skills. These models are trained with reinforcement learning, where the model is rewarded for being correct and penalized when incorrect. This training system uses deep learning to help models develop reasoning skills across a range of domains.

While reasoning models are generally more powerful than chain of thought, it's often worth testing both approaches for your use case. Thinking models strive to reason about everything effectively, but a well-crafted chain of thought prompt from a human expert can often outperform them when developing use-case-specific models/APIs.

</details>

### How Kiln handles reasoning models and chain of thought

Kiln has native support for both these methods. This includes:

* [**Creating Reasoning Models**](#building-your-own-reasoning-model-distillation)**:** You can fine-tune/distill reasoning models using Kiln. Models train on your Kiln dataset, using samples generated from reasoning models. This approach allows you to build small, fast, and high-quality thinking models, tuned to your use case.
* [**Data Model**](/developers/kiln-datamodel)**:** Our data model stores thinking separately from final answers, allowing you to evaluate or train on them independently.
* [**Custom Message Flow**](#custom-message-chat-flow)**:** When using chain-of-thought with models that don't support reasoning we make a chain of calls to the model to formally separate the thinking from the answer.
* [**Structured Data**](/docs/structured-data-json): Our chat call flow allows for the final answer messages to use structured data tools (json\_schema, json\_object, tool-calls, etc.) without adding "thinking" fields to your data structures.
* [**Prompts**](/docs/prompts)**:** Prompts are divided into the primary system message and a separate "thinking instruction."
* **Reasoning Parsers:** Kiln includes parsers that separate out "thinking" from answers for common thinking models.

### Using Reasoning or COT in Kiln for inference

Using reasoning models or COT in Kiln is easy! Simply do one of the following:

* Run a model with any "Chain of Thought" prompt selected, including a custom prompt with thinking instructions included.
* Run a reasoning model (e.g., Deepseek R1) with any prompt. If a Chain of Thought prompt is used, the thinking instructions will be passed along to the model in the system prompt. If a non-COT prompt is used, the reasoning model will still "think," but using its own reasoning guidance.

Once the run is complete, you'll see both a final answer and reasoning in the model output.

### Building your own reasoning model (distillation)

Kiln can fine-tune a thinking model using your dataset. Often called distillation, these models can learn the reasoning strategies for your use case from examples in your Kiln dataset. By fine-tuning a model, you can produce a model that's smaller, faster, cheaper, and better than the original model (for your use case).

* See our guide on fine-tuning reasoning models
* See our guide for [general fine-tuning](/docs/fine-tuning/fine-tuning-guide) (non-reasoning models)

{% hint style="info" %}
Kiln uses supervised fine tuning to distill reasoning models.
{% endhint %}

### Performance & Cost

Reasoning and COT doesn't necessarily mean slower or more costly requests. Sometimes a smaller model with these methods can be faster, better, and cheaper than a larger model performing the same task. Fine-tuning can help further reduce costs and improve quality.

### Supported Reasoning Models

Currently we support Deepseek R1 and its official distillations. We expect to see many more open reasoning models emerge over the next few months. See the [model capability list in our docs](/docs/models-and-ai-providers#included-models-recommended) for the latest .

{% hint style="info" %}
While you can call OpenAI's reasoning models (o1, o3) from Kiln, they behave like normal models. OpenAI hides the reasoning tokens from users, only returning the final answer.
{% endhint %}

### Custom Message Chat Flow

Here's the message call flow Kiln uses for each configuration:

#### Normal Call-Flow (non-reasoning model)

* \[System-Message]: System prompt
* \[User-Message]: User inputs
* \[Assistant-Message]: Final Answer, optionally structured data

#### Chain of Thought Call-Flow (non-reasoning model):

* \[System-Message]: System prompt
* \[User-Message]: User inputs
* \[User-Message]: Thinking instructions. User-provided if available, defaults to "Think step by step, explaining your reasoning."
* \[Assistant-Message]: COT reasoning tokens
* \[User-Message]: Kiln managed message: "Considering the above, return a final result."
* \[Assistant-Message]: Final Answer, optionally structured data

This flow is also used on fine-tunes you create with Kiln if the fine-tune was created with the "Final answer and intermediate reasoning" training strategy.

#### Reasoning Model Call-Flow:

* \[System-Message]: System prompt, optionally appending thinking instructions if the selected prompt includes them.
* \[User-Message]: User inputs
* \[Assistant-Message]: Final Answer and reasoning in one message, but will be parsed into separate reasoning and answer fields. Will parse structured data if the task has structured output.


# Reviewing and Rating

Ratings help multi-shot prompting, fine-tuning, evals, and more

Kiln includes a rating interface for rating dataset entries. This can be used to score the quality of the generated data, or to evaluate the quality of a model.

<figure><img src="/files/VwWMsBpDNPIKxUHhGWfZ" alt="" width="341"><figcaption><p>Rating UI in Kiln Desktop</p></figcaption></figure>

### When Rating Options Appear

You'll see rating options on your dataset items:

* The "Overall Rating" will option will always appear
* After creating an [Eval](/docs/evals-and-specs/evaluations), rating options will be visible for each sample in its golden dataset.

Not every rating will appear on every data sample, and that's okay! They only appear when they are useful, such as aligning a LLM-as-judge to human preference with a golden dataset. If a specific rating doesn't appear, it means it wouldn't be used and isn't necessary to rate this item by that criteria.

Want to rate an item that isn't showing a rating field? Add a [tag](/docs/organizing-datasets#using-tags-to-organize-your-dataset) like "eval\_NAME\_golden" which tells the system how that rating should be used. Once tagged, the necessary ratings will appear.

{% hint style="info" %}

#### Legacy Task Requirements

Older version of Kiln had the concept of "task requirements": a rating criteria for all dataset samples. We've removed these going forward. Rating every single data sample by a criteria isn't necessary or helpful. As described above we now show the right ratings on the right items, and nothing more.
{% endhint %}

### Rating Option Parameters

Each rating option has a number of parameters:

* Name: the name of the requirement, which will appear in the rating UI. Limited in length to fit in the UI, but you can add more content in the instructions field below.
* Instructions: more details about the requirement. These will be available to reviewers in the UI (under the ![](/files/lRsaGzU4HVSstGzQ8ZEv) icon).
* Rating Type: one of 5-star, pass/fail, pass/fail/critical.
* Priority: how important this criteria is to the task.

### Rating Types:

* **5-star**: a 1-5 star rating.
* **Pass/Fail**: A binary pass/fail rating.
* **Pass/Fail/Critical**: A ternary pass/fail/critical rating. It can be useful to add the "critical" level when there are criteria where some failures are exceptionally important to avoid. For example, a customer service bot could have a "tone" criteria, where casual/slang language would be a failure, but profanity or insulting the user would be critical.
* **Custom**: you can define a custom rating scale when using python library. However, you won't be able to use custom ratings in the Kiln UI.

### How Ratings are Used

Kiln uses ratings in a variety of ways:

* In evals, ratings of your golden dataset are used to benchmark and compare judges for evaluating your task. This helps you find the [ideal judge](/docs/evals-and-specs/evaluations#finding-the-ideal-judge).
* Kiln's [automatic prompt generators](/docs/prompts#prompt-generators) may incorporate highly rated samples into a prompt as a few-shot example. These filters to examples 4+ stars, and prefers 5-star ratings if available.
* When creating a [fine-tuning dataset](/docs/fine-tuning/fine-tuning-guide), you may optionally filter the training data to highly rated content.
* When using the [python library](/developers/python-library-quickstart), you can access or set ratings.


# Collaboration

How to collaborate with your team using Kiln

<figure><img src="/files/DcJ2jIvEEPfUciIWh13u" alt=""><figcaption></figcaption></figure>

### Designed for Techies and Non-techies Alike

It's easy to collaborate with Kiln across teams with technical team members (devs, data-scientists), and non-technical team members (subject-matter experts, QA, labelers, etc).

We suggest [Git](#option-1-use-git) for technical teams, [shared drives](#option-2-use-shared-drives-for-non-technical-team-members) for non-technical teams, or a [mix](#option-3-combining-git-and-shared-drives) for mixed teams.

### **Recommended: Use Git!**

Kiln projects are simply a folder of files, making it easy to share them using Git. Add your project folder(s) to a git repo and you're set up with an excellent collaboration workflow with branches, pull requests, version control, access control, and more!

Kiln's [**Automatic Git Sync**](/docs/collaboration/automatic-git-sync) feature makes it easy for non-technical team members to collaborate, even if they have no idea what git is. It's the best way to share Kiln projects with less technical team members like subject matter experts, QA, or PMs.

See [below](#collaboration-design) for how our file format is optimized for Git-based workflows.

### Option 2: Use Shared Drives

You can also host a Kiln project on a shared drive of your choice (Google Drive, Dropbox, iCloud, etc).

Be sure to review our [**Automatic Git Sync**](/docs/collaboration/automatic-git-sync) feature before choosing a shared drive. It's more powerful, easier to setup, and harder to break than a shared drive.

Kiln project files will track who created them (internally in their JSON), which adds version history when many folks are making changes on the same shared drive. It's not as robust as Git history, but there is attribution built in.

### Collaboration Design

As you may have already guessed, you don't need to allow a third party to access your data, or maintain a database. Everything runs locally on your machine, and syncs through existing tools you control.

Kiln's data structure was designed with collaboration in mind:

* A Kiln project is simply a folder of files, which makes it compatible with a range of existing collaboration tools, from Git to Dropbox
* New items use unique random IDs to avoid conflicts/collisions, allowing many people to work concurrently on the same project.
* Project files are kept small and predominantly append-only. It's rare multiple people will need to work on the same file at the same time, reducing conflicts.
* The Kiln project files are JSON files, and are formatted to be easily used with diff tools and standard PR tools (GitHub, GitLab, etc).
* Static paths: even when changing the name of resource, the path will remain static.

### Kiln is an App: Don't Deploy Kiln as a Service

{% hint style="info" %}
This section is for engineers/developers attempting custom deployments. If you're a normal app user who launches Kiln as an app, just skip this!
{% endhint %}

Kiln is a desktop app, designed as an app. Even though internally it uses web tech (HTML), it's still an app designed to be run locally on each user's machine, not a hosted service.

We don't recommend or support trying to host it as a service and access it over the network. There are several major downsides/risks if you do.

<details>

<summary>Technical Details</summary>

There are several downsides/risks if you try to run as service:

* Security: there's no web-based logins or access controls, so anyone who can access the service can edit data and send requests. That's okay when running as an app locally behind your machine login, but brings risk when opening the service up to anyone over a network.
* Collaboration: If multiple users are sharing an instance of Kiln, all the created\_by tags in your dataset will all have the machine name of your VM/host, not the individuals.
* Data backup and history: if you run as suggested above, Git and/or the shared drive will provide data sync, backup and history. If you run on a single server, there's a higher risk of data loss if that server drive is lost/damaged.
* Native system Integrations: Kiln accesses OS features like the taskbar, local filesystem and custom folder-picker UI; these aren't available to web apps. We continue to add more native integrations over time.

The solution is just to have each user run Kiln locally on their own machine and collaborate using the two mechanisms above. It's quite easy and there are no compromises:

* We support all major platforms: Mac, Windows and Linux. Each has a user-friendly [installer](https://kiln.tech/download).
* You don't need to co-locate Kiln and your LLM/GPU servers. You can still connect Kiln to LLM services on another machine by setting a custom URL. See [AI Provider Setup](/docs/models-and-ai-providers) for details.
* Kiln doesn't require a GPU and uses minimal system resources (CPU, memory).
* Kiln can work offline, and instantly. No network lag.

</details>


# Automatic Git Sync

Git, even if you've never used a terminal

<figure><img src="/files/Ze5jCwtgEGpxW2VB49zr" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
**Read our** [**technical deep dive on Kiln's automatic git sync**](https://kiln.tech/blog/git_backed_saas)
{% endhint %}

Historically, there have been two methods of collaboration for data teams: Git or cloud services. Both had tradeoffs.

**Kiln adds a new option: automatic Git sync — providing the best of both worlds for data teams.**

|                                    | Git                                                                       | Cloud Service                      | Kiln Automatic Git                                    |
| ---------------------------------- | ------------------------------------------------------------------------- | ---------------------------------- | ----------------------------------------------------- |
| **Ease of Use**                    | ❌ Requires terminal commands, dev tools, manual sync, conflict resolution | :white\_check\_mark: Log in and go | :white\_check\_mark: Connect and go                   |
| **History / Traceability**         | :white\_check\_mark: Detailed, traceable, revertible                      | ❌ Limited                          | :white\_check\_mark: Detailed, traceable, revertible  |
| **Data Portability**               | :white\_check\_mark: Easy portability on open formats                     | ❌ Custom formats, low portability  | :white\_check\_mark: Easy portability on open formats |
| **Privacy / Enterprise Approvals** | :white\_check\_mark: Existing Git host in 99% of orgs                     | ❌ Requires custom approvals        | :white\_check\_mark: Existing Git host in 99% of orgs |

### How it Works

Kiln allows you to choose how you work:

* **PMs / Subject Matter Experts / Designers:** connect an existing Kiln project in a repository in seconds via Kiln's UI. Any work you do (ratings, evals, data annotation) is automatically and instantly synced to Git, without any manual intervention. They never see a terminal, or have to learn Git.
* **Developers**: use manual Git tools, or automatically sync, or a mix. It's up to you!

#### Step 1: Set up your project repository

This is the only step requiring a developer or someone familiar with Git. All other users can use Kiln via the UI, with no knowledge of Git!

To set up your project:

* Create a Kiln project using the Kiln UI (or choose an existing project)
* Create an empty remote repository on GitHub, or the provider of your choice
* Push the new project to your remote repository. Follow your provider's instructions, but typically something like `git init -b main && git add . && git commit -m "initial commit" && git remote add origin <your-repo-url> && git push -u origin main`
* Ensure your team has read/write access to the repository using the provider's user management

{% hint style="info" %}
**Picking a Git Host**

Kiln works with any remote Git host. Most orgs already have a preferred provider, but there are small differences between how they work with Kiln.

* **GitHub**: Just a few clicks to install via OAuth. It's easiest if the dev who initially sets up the project also installs and authorizes the app once, which will skip the "install app" step for all future users.
* **GitLab**: one-click link to generate a personal access token. Easy even for non-technical users. May require approvals based on organization policies.
* **Others**: supported via "personal access tokens". You'll need to provide your less technical users instructions for how to get a token for your chosen host.

Once connected: all hosts behave the same. The differences are only during setup.
{% endhint %}

#### Step 2: Share a "Connect" link with your Team

* Navigate to **Settings > Manage Projects > Import Project > Automatic Git Sync** in the Kiln UI
* Paste the https URL of your repo. For example `https://github.com/Kiln-AI/sync_test` or `https://github.com/Kiln-AI/sync_test.git`
* Copy the URL from the browser — this is now a deep link to add this project on any team member's machine.
* Share the link with your team! If you're using a provider other than GitHub or GitLab, also share instructions on how to generate a personal access token.

#### Step 3: Connect the Repository

New users can simply:

* Open the connect link generated above
* Follow the in-app steps to connect the project
* Use it like any other project, except know it's automatically kept up to date with your team!

{% hint style="info" %}
This path can be used for any user, regardless of technical background. The cloning, syncing, authorization, and all other details are managed automatically.
{% endhint %}

### What about Conflicts / Rebasing / etc?

**Short answer:** if you're using Kiln's automatic sync you don't need to worry about it.

Kiln has a powerful sync engine under the hood.

* Each project is cloned to its own private and hidden repo to avoid conflicts
* The Kiln data model was already designed for Git: small immutable files, typically append only. The chance of conflicts is kept extremely low.
* We keep your local repo up to date with your remote repository (within 15 seconds). There's almost no time for drift or conflicts.
* If conflicts do occur (rare for above reasons), Kiln is self healing. It immediately detects the issue, rolls back and stashes any changes (zero data loss), and gets back into sync with remote. The UI shows a clear error; you can just hit "Submit" again and your change should now go through.

### Repository Structure

You can put the Kiln project in any folder of your repository. We'll scan for `project.kiln` files on first sync, and let you select which project to use if multiple exist.

This also means you can put many Kiln projects into a single repository (don't worry, they won't cause conflicts).

### SSH Access

Kiln will work with SSH access URLs if your machine already has SSH set up with your Git provider. Just paste the SSH URL starting with `git@` when setting up the repository.

However, non-devs are very unlikely to have SSH set up. If SSH fails, we'll suggest trying the https URL and OAuth/token-based authentication instead.


# Organizing Datasets

Tag, Filter, Sort, Import, Freeze and Split

<figure><img src="/files/DhQlK1M86PnkX4yN2IRW" alt=""><figcaption></figcaption></figure>

### Using Tags to Organize Your Dataset

Kiln uses tags to organize your dataset. You can add tags to any run/sample, and then filter by tag. This is a great way to organize your dataset and find specific runs.

Some examples of how you might use tags within a team:

* Working with eval teams:
  * Add the "needs\_review" tag when data is ready for review by a human eval team
  * Ask a human eval team to review the new batch of synthetic data. New synth data is automatically tagged with "synthetic" and "synthetic\_session\_id".
* Defining a "golden" dataset: Have QA tag a "golden" data reserved for evals
* Bug resolution: QA can tag examples of a common issue with a tag (e.g. "issue\_unprofessional\_tone"). Data scientists can run evals of different methods of fixing the issue.
* Regression Testing: tag important customer use cases with a tag ("customer\_use\_case") and run evals to ensure the model doesn't regress on them prior to a new release.
* Fine-tuning: exclude tags from fine-tuning datasets (golden, customer\_use\_case, etc), to prevent contamination.

### Sort and Filter

The dataset view offers a number of tools that make working with large datasets easier:

* Filter: Tap the filter button (![](/files/1pjnOvrNQjzEfF8caV7V)) to filter to specific tags
* Sort: You can sort by any column by clicking its header
* Multi-select: you can enter "selection" mode by clicking the select button
  * Select any row by clicking it
  * Select a range of rows by clicking the first, then holding shift while clicking the last

### Batch Editing

Once you have selected rows you can perform a number of batch actions:

* Add tags
* Remove tags
* Delete dataset items

### Importing Data into your Dataset

If you already have a dataset, it's easy to import it into Kiln. Open the dataset tab, then click "Upload File" to add your data.

The format must be a CSV file with a header row. The following columns are supported:

* `input` \[Required] - The input to the task. If the task has an input schema, this must be a JSON string conforming to that schema.
* `output` \[Required] - The output of the task. If the task has an output schema, this must be a JSON string conforming to that schema.
* `reasoning` \[Optional] - If you model is a reasoning model that output reasoning/thinking text before the output (for example, R1, QwQ, etc), you can provide that text here. This will be visible in the UI, and available for fine-tuning a reasoning model.
* `chain_of_thought` \[Optional] - If you model output chain-of-thought text before the output, you can provide that text here. This will be visible in the UI, and available for fine-tuning a thinking model.
* `tags` \[Optional] - comma separated string listing the tags you want to add to this row. For example: `tag1, tag2`.

<figure><img src="/files/GE7hYHWz3Ebr3emPfPrA" alt="" width="375"><figcaption><p>The CSV Import UI</p></figcaption></figure>

{% hint style="info" %}
If you prefer working in python, or have a complex import use case, our Python SDK can be used to add data to a Kiln project. It includes validators that ensures your data conforms to the needed schemas.

See our [python docs for an example](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#load-an-existing-dataset-into-a-kiln-task-dataset).
{% endhint %}

### Dataset Splits

When creating a fine-tune, you can define a "dataset split". This is a frozen subset of your data.

* Dataset splits may be broken into sub-sets like "train", "validation" and "test" which are useful for systematically training and evaluating models.
* Dataset splits will randomly assign items between sub-sets (train/test/val), but the assignment is static. Items do not shift between subsets once the dataset split is created.
* Dataset splits do not grow/change when you add new data. They are frozen at the point in time when they are created. This makes it easier to run multiple experiments (fine-tunes, evals, etc) on exactly the same training/eval datasets.


# Structured Data / JSON

<figure><img src="/files/UEFYsLdVWk3NAdntmM63" alt=""><figcaption></figcaption></figure>

Structured data is a first-class citizen in Kiln.

* JSON input/output is supported for any task.
* You can define input/output schemas for each task you create.
* We automatically detect when a AI model doesn't produce output in the correct format.
* No data will be saved into the dataset without first passing validation, which keeps the dataset clean.
* Our [included models](/docs/models-and-ai-providers#included-models-recommended) are tested for JSON compatibility, and models that don't perform well with structured data will show a warning if you attempt to use them on tasks with structured output.

### Creating a Schema

Our easy-to-use visual schema builder lets you create and use structured schemas. Define your input and output schemas in our UI when creating a task.

{% hint style="info" %}
You can't edit the input/output schemas after creating a task, as that would invalidate all prior data. Create a new task with your updated schema instead.
{% endhint %}

<figure><img src="/files/5kwPiHFLuaW2x6YcBCLa" alt="" width="375"><figcaption><p>Kiln's Visual Schema Editor</p></figcaption></figure>

### Raw JSON Schema

For technical users, we support any valid [JSON-schema](https://json-schema.org) for inputs and output schemas.

JSON schema is more powerful than the visual editor, allowing arrays, nested objects, enums, constraints and more. You can use raw JSON schemas from the Python library, or from the UI (by picking an array type).

{% hint style="warning" %}
We only recommend JSON schema for technical users/developers. It's much more complex than the visual editor.

Errors in the schema will likely result in bugs when running your task.
{% endhint %}

### Troubleshooting Structured Data Issues

LLMs aren't always great at producing valid JSON output. Kiln does its best to use a variety of techniques to get the correct output, but it's never guaranteed.

You may see errors in the UI if the model produces invalid JSON, or JSON that doesn't match the output schema of your task.

Here are some tips to getting consistent JSON output:

* Use a model that is suggested for JSON output. The model drop down will warn you if you're using an untested model, or a model with known issues producing JSON.
* Consider [fine-tuning a model](/docs/fine-tuning/fine-tuning-guide) on synthetic data from a larger and more reliable model. Fine tuning is a great way to get consistent format output, and you can get models as small as 1B parameters to consistently produce JSON with fine-tuning.
* Don't use an output schema unless necessary. If you only require plain text output from your task, be sure to select plaintext as the output format when creating the task. This will work with well with any model.
* Try adding an example of the output structure you want to your prompt. You can do this in `Settings > Edit Task > Task Instructions` or via a [custom prompt](/docs/prompts#custom-prompts-saved-prompts). For example, add this to your prompt:

````
Provide the output in the following JSON format:
```
{
  "setup": "Why did the chicken cross the road",
  "punchline": "To get to the other side."
}
 ``` 
````


# Keyboard Shortcuts

<figure><img src="/files/ldPEP6qxImTGKllzVahL" alt=""><figcaption></figcaption></figure>

Kiln offers a number of keyboard shortcuts which makes navigating the app easier:

* After opening a dataset item from the Dataset screen, use right and left keys to navigate to the next and prior item.
* When focusing a 5-star rating, using the 1-5 keys to select a value, or 0 to clear the value
* When viewing a dataset item, use the delete key to delete that item.
* When on the Run screen, you can invoke a run with Ctrl+Enter on Windows, or Command+Return on macOS.
* Elsewhere: wherever you see a keyboard shortcut preview on a button, such as ![](/files/OYSspAHnGAoY0kCzv8tD), you can enter that shortcut to activate the button where the preview appears.
* See the [organizing datasets](/docs/organizing-datasets) docs page for how to multi-select dataset items


# Privacy

Kiln is private by design

<figure><img src="/files/ma7bPMFmcerw2I4kcySt" alt=""><figcaption></figcaption></figure>

Kiln is local software that runs on your computer, not the cloud. Your dataset is stored locally on your own hard drive.

We don't have access to your dataset, and couldn't access it even if we wanted to.

All of Kiln's source code is on GitHub. Anyone can verify these privacy claims. Kiln's binary builds are built from the public source code, on public CI servers (Github Actions), and have verifiable checksums.

Requests to third party AI providers (like OpenAI or OpenRouter) are sent directly from your computer to those API providers. We can't see your keys or the data from those requests.

#### Desktop App Analytics

The Kiln desktop app collects analytics so we can understand how people are using it. This includes analytics like which pages are being visited, and which actions are taken in the app’s UI. Analytics are collected with [Posthog](https://posthog.com/).<br>

* The Kiln Python library does not collect analytics. Only the desktop app’s user interface collects analytics.
* Analytics never include your dataset information, such as inputs to models, output from models, project names, etc.
* Analytics never include your API keys.
* Analytics may be associated with your email address if you register your email address.

{% hint style="info" %}
If you opt-in to our email newsletter, we collect your email address to provide you the newsletter. You can unsubscribe any time.
{% endhint %}


# Repairing Responses

"Teach the model, you will" - ML Yoda

<figure><img src="/files/ZfW2yft636OK3RBNDfSu" alt=""><figcaption></figcaption></figure>

Kiln offers a somewhat novel option in our "Run" UI and data model: repairs. Instead of simply evaluating/rating AI responses, we offer the ability to "repair" any responses which are reviewed less than 5-stars.

### Repair Process

The repair process is:

1. Rate the model response. If the overall rating is 5-star, you're done!
2. A human offers repair instructions, describing what was wrong, and how it can be improved, then click "Repair".
3. The model generates a new response, given the prompt, initial flawed response, and the human repair instructions.
4. A human evaluates the response, and either accepts it as 5-star (you're done!), or returns to step 2 to iterate on repair instructions until the model can produce a 5-star response.

### Why Repair Instead of Simply Correct the Response

Repairing generates some useful data in our dataset, such as:

* Examples of failures.
* Feedback on failures. Specifically, feedback we know the model can understand, because it has demonstrated it can generate a 5-star response after receiving it.
* More nuanced examples: the difference between a 4-star and 5-star response for the same input can be helpful for model evaluation and training.

### Where is this data used?

Currently this data is used in [repair prompts](/docs/prompts#prompt-builders-prompt-styles): a prompt style that includes both all of the data above in a multi-shot prompt format.

In the future we plan to use it in other places, like evaluations, so collecting it now will help you in the future.

You can also use this data for your own evaluation processes via our [python library](/developers/python-library-quickstart). For example, clustering repair feedback to identify recurring issues.

### Why can't I just manually correct errors?

This is planned! For now, please use the repair system for corrections.


# Troubleshooting & Logs

How to troubleshoot errors from the logs

### First: Consider Model Errors

Most bug reports we get aren't actually bugs in Kiln, but models failing to perform the task asked of them. Please read the error messages carefully, and follow the suggestions.

* If the error is related to structured data, see [Troubleshooting Structured Data Issues](/docs/structured-data-json#troubleshooting-structured-data-issues)
* Try another (smarter) model to see if the issue persists, of if it's isolated to a specific model.

### Second: Review the Error Message For Guidance

Often the error message explains how to resolve the issue. Some examples:

* Issues from connected services (no credit, rate-limits, etc) will need to be corrected with those services
* "Model not found, inaccessible, and/or not deployed" - A Fireworks API message that the model needs to be re-deployed. For fine-tunes, Kiln can re-deploy the model for you - just open the fine-tune in the "Fine Tune" tab.

### Third: Check The Logs

Kiln writes out error logs to the path \`\~/.kiln\_ai/logs/\*.log\`.

Check your latest log file for a stack trace of any errors, which can help understand the issue. If you think it's a bug in Kiln and not a model issue, please [file a bug](https://github.com/Kiln-AI/Kiln/issues), following the bug template and attaching logs.


# Productionizing Kiln

How to take ideas from the lab (Kiln) into your product

Kiln is a great place to rapidly experiment with many models, providers, fine-tunes, and prompts.

Once you've found the ideal way to run your AI workload, you are ready to create (or iterate) on your product. This guide walks through the options for creating products from Kiln projects.

### Productionizing Kiln Tasks

#### Option 1: Use the Kiln Python Library

The MIT open-source [Kiln python library](/developers/python-library-quickstart) powers all of the requests in the Kiln app. You can use it in your app or server, pointing it to the Kiln project of your choice.

Our python library docs have instructions for [exporting your project and running it via python](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#run-a-kiln-task-from-python).

#### Option 2: Exporting Kiln Search Tools / RAG

We have multiple options for deploying search tools you build inside of Kiln. See the docs for options for [deploying search tools / RAG](/docs/documents-and-search-rag#deploying-your-rag).

#### Option 3: \[Simple Tasks Only] Replicate Kiln Prompts and Requests in the Framework of your Choice

{% hint style="warning" %}
This approach works well for simple single-turn AI tasks. However replicating a complete agent system with sub-agents, tools, or RAG can be quite challenging. We suggest using the Kiln library for complex tasks.
{% endhint %}

Every model request sent from Kiln is logged to `~/.kiln_ai/logs/model_calls.log` . This includes the messages and any parameters sent to the API (model, temperature, reasoning, etc). It even includes a curl command to replicate the call Kiln made from the command line.

You can use the data in this log to exactly replicate any Kiln request in the language/framework of your choice. You can choose which tech-stack makes sense for your product; anything from simple HTTP requests, to integrating provider libraries (`pip install openai`), to integrating [llama.cpp](https://github.com/ggml-org/llama.cpp) into your app/process.


# Contact Us

We'd love to chat!

Feel free to reach out:

* Bugs and Feature Requests: [Github Issues](https://github.com/Kiln-AI/Kiln/issues)
* Chat: [Discord](https://kiln.tech/discord)
* Newsletter: [Newsletter Signup](https://kiln.tech/blog)
* Blog: [Kiln Blog](https://kiln.tech/blog)
* Ideas & Discussions: [Github Discussions](https://github.com/Kiln-AI/Kiln/discussions?discussions_q=)
* Email: support at our\_domain
* Security Reports: [Github Security](https://github.com/Kiln-AI/Kiln/security)
* [Enterprise Sales](https://kiln.tech/contact_sales)


# Python Library Setup

pip install kiln\_ai

{% hint style="info" %}
The Python library is completely optional. If you just want to use the app, follow our [Quickstart guide](/docs/quickstart).
{% endhint %}

Our open source [python library](https://pypi.org/project/kiln-ai/) allows you to use Kiln from your codebase. This can be as simple as accessing a Kiln dataset from a notebook, or as advanced as running any of our features from code.

[![PyPI - Version](https://img.shields.io/pypi/v/kiln-ai.svg?logo=pypi\&label=PyPI\&logoColor=gold)](https://pypi.org/project/kiln-ai/) [![PyPI - Python Version](https://img.shields.io/pypi/pyversions/kiln-ai.svg?logo=python\&label=Python\&logoColor=gold)](https://pypi.org/project/kiln-ai/) [![Docs](https://img.shields.io/badge/docs-pdoc-blue)](https://kiln-ai.github.io/Kiln/kiln_core_docs/index.html)

### Installation

To install Kiln, run the following command:

```bash
pip install kiln_ai
```

### Library Docs & Examples

Our [library docs](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html) include an API reference for all of the library's features.

The docs include several quick-start examples to get up and running:

* [Using the Kiln Data Model](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#using-the-kiln-data-model)
  * [Understanding the Kiln Data Model](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#understanding-the-kiln-data-model)
  * [Datamodel Overview](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#datamodel-overview)
  * [Load a Project](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#load-a-project)
  * [Load an Existing Dataset into a Kiln Task Dataset](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#load-an-existing-dataset-into-a-kiln-task-dataset)
  * [Using your Kiln Dataset in a Notebook or Project](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#using-your-kiln-dataset-in-a-notebook-or-project)
  * [Using Kiln Dataset in Pandas](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#using-kiln-dataset-in-pandas)

### Web/REST API

We also offer a self-hostable [REST API](/developers/rest-api) for Kiln, based on FastAPI.


# Rest API

Call Kiln over HTTP

[![PyPI - Version](https://img.shields.io/pypi/v/kiln-server.svg?logo=pypi\&label=PyPI\&logoColor=gold)](https://pypi.org/project/kiln-server/) [![Docs](https://img.shields.io/badge/docs-OpenAPI-blue)](https://kiln-ai.github.io/Kiln/kiln_server_openapi_docs/index.html)

We also offer an open-source self-hostable REST API (web API) for Kiln, based on [FastAPI](https://fastapi.tiangolo.com).

To install the Kiln rest API, run the following command:

```bash
pip install kiln_server
```

### Docs

The [REST API docs](https://kiln-ai.github.io/Kiln/kiln_server_openapi_docs/index.html) explain the endpoints, parameters, response format, and errors.

### Client Side Libraries

The REST API supports OpenAPI, which means you can generate typed client libraries for almost any language using code generators:

* [OpenAPI Generator List](https://openapi-generator.tech/docs/generators/)
* [Swagger Codegen](https://github.com/swagger-api/swagger-codegen)


# Kiln Data Model

How Kiln projects are structured

### Understanding the Kiln Data Model

Kiln projects are simply a directory of files (mostly JSON files with the extension `.kiln`) that describe your project, including tasks, runs, ratings, fine-tunes and other data.

This dataset design was chosen for several reasons:

* Git compatibility: Kiln project folders are easy to collaborate on with Git (or a shared drive). See our [collaboration guide](/docs/collaboration#technical-collaboration-architecture) for additional details of how we avoid conflicts and format to support diff tools.
* JSON allows you to easily load and manipulate the data using standard tools (pandas, polars, etc.).

### Data Model Overview

Here's a high level overview of the Kiln datamodel. A project folder will reflect this nested structure:

* Project: a Kiln Project that contains related tasks.
  * Task: a specific task including prompt instructions, input/output schemas, and requirements.
    * TaskRun: a sample (run) of a task including input, output, and human rating information.
    * [Finetune](/docs/fine-tuning/fine-tuning-guide): a model for fine-tuning jobs. Includes configuration, status tracking, and data necessary to call the deployed fine-tuned model.
    * [Prompts](/docs/prompts): a custom prompt for this task. See our [prompts docs](/docs/prompts) for details.
    * DatasetSplit: a frozen collection of task runs divided into train/test/validation splits.
    * [Task Run Config](/docs/evals-and-specs/evaluations#finding-the-ideal-run-method): a specific method of running this task, including a prompt and model.
    * [Eval](/docs/evals-and-specs/evaluations): an evaluation of this task, including output score definitions and datasets to run this eval on
      * [Judge](/docs/evals-and-specs/evaluations#finding-the-ideal-judge): a method of running this eval, including instructions and model.
      * Eval Runs: the results of running this eval, including scoring.

See the [python library datamodel docs](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai/datamodel.html) for detailed descriptions of classes, fields and validations.

### Python Library

If you want to access the data model via code, check out our [python library](/developers/python-library-quickstart). The library offers iterators, typed classes, pydantic validation, and more. It's the easiest way to read and mutate a Kiln dataset.

### Direct Access

You can load Kiln project files using any tool which supports JSON, including polars and pandas. See the [example](https://kiln-ai.github.io/Kiln/kiln_core_docs/kiln_ai.html#using-kiln-dataset-in-pandas) in our library docs.

{% hint style="warning" %}
We highly recommend the Kiln python library for any writes to `.kiln` files. It will run validators which catch issues which could break your project.
{% endhint %}


