Golden Datasets and CI Regression Tests for LLM Apps in October 2026: Catch Prompt and Model Changes Before They Ship
Most teams test their code and then change a prompt, swap a model, or tweak a retrieval setting with nothing but a quick manual try. The change looks fine on three examples, ships, and quietly breaks a fourth case nobody checked. This post shows a practical way to stop that: a small golden dataset, a handful of cheap automated checks, and a CI job that fails when quality drops.

What a golden dataset is (and what it is not)
A golden dataset is a fixed set of inputs, each paired with a description of what a good output must satisfy. It is not a benchmark and it does not need thousands of rows. For most apps, 30 to 100 well-chosen cases catch the majority of regressions, because the point is coverage of your own failure modes, not statistical precision.
Good sources for cases, roughly in order of value:
- Real failures. Every bug report or bad answer you fix becomes a test case, so the same mistake cannot return silently.
- Real traffic, sampled and scrubbed. Pull representative requests from logs, remove personal data, and keep a mix of easy, typical, and awkward ones.
- Edge cases you invent. Empty input, very long input, mixed languages, requests the model should refuse, and requests that look similar but need different handling.
Store the dataset in version control next to the prompts, as JSONL or YAML. Reviewers should see a dataset change in the same pull request as the prompt change that motivated it.
Step 1: Define the case format
Keep each case small and explicit. A useful shape looks like this:
{"id": "refund-policy-01",
"input": {"question": "Can I return an opened item after 20 days?"},
"checks": {
"must_contain": ["30 days"],
"must_not_contain": ["guarantee"],
"json_schema": "answer_v1",
"max_chars": 600
},
"tags": ["policy", "core"]}
The id makes failures easy to find, checks hold the rules, and tags let you run a quick subset on every commit and the full set nightly.
Step 2: Prefer deterministic checks first
Model-graded scoring is useful, but it is slower, costs money, and adds its own noise. Start with checks that are plain code:
- Schema validity. If the output should be JSON, parse it and validate it against a schema. This alone catches a large share of prompt and model-change breakage.
- Required and forbidden strings. Check that key facts appear and that banned phrases do not.
- Length and format limits. Maximum length, required headings, no leftover template markers.
- Tool-call expectations. For agents, assert which tool was called and with which arguments, not the wording of the reply.
- Refusal behavior. For inputs that should be declined, check for the refusal pattern you expect.
Reserve a model-based grader for the things code cannot judge, such as whether a summary is faithful to its source. If you use one, give it a narrow rubric, ask for a pass or fail with a short reason, and spot-check its verdicts by hand now and then.
Step 3: Write the runner
A runner loads the cases, calls your real application code (not a re-implementation of it), and applies the checks. Here is a minimal pytest version in Python:
import json, pathlib, pytest
from jsonschema import validate
from myapp.llm import answer # your real entry point
CASES = [json.loads(l) for l in
pathlib.Path("evals/golden.jsonl").read_text().splitlines() if l.strip()]
SCHEMAS = {"answer_v1": json.loads(pathlib.Path("evals/answer_v1.json").read_text())}
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
def test_golden(case):
out = answer(**case["input"]) # returns a string
checks = case["checks"]
for s in checks.get("must_contain", []):
assert s.lower() in out.lower(), f"missing: {s}"
for s in checks.get("must_not_contain", []):
assert s.lower() not in out.lower(), f"forbidden: {s}"
if "max_chars" in checks:
assert len(out) <= checks["max_chars"]
if "json_schema" in checks:
validate(json.loads(out), SCHEMAS[checks["json_schema"]])
Using parametrize means each case is its own test, so a failure report names the exact case that broke instead of one big red result.
Step 4: Handle nondeterminism honestly
Model outputs vary between runs, which is the reason many teams skip automated tests. You can manage it without pretending it away:
- Lower the temperature for the test run if your provider supports it, and set a seed where one is offered. Treat seeds as best effort, since not every provider guarantees identical output.
- Test properties, not exact text. The checks above assert facts, structure, and limits, which survive harmless wording changes.
- Retry flaky cases a fixed number of times and record the pass rate. A case that passes 2 of 3 runs is a signal worth looking at, not noise to hide.
- Use thresholds for the whole suite. For example, require every case tagged
coreto pass, and allow a small agreed number of failures in the long tail while you fix them.

Step 5: Wire it into CI without burning money
Running real model calls on every commit can get expensive and slow. Split the suite into tiers:
- Every pull request: the
coretag only, plus all deterministic checks. Keep it to a few minutes. - Pull requests that touch prompts, model identifiers, or retrieval settings: the full suite. You can trigger this with a path filter on the folders that hold those files.
- Nightly: the full suite plus any model-graded checks, with results stored so you can chart trends.
A minimal GitHub Actions job for the second tier looks like this:
name: llm-evals
on:
pull_request:
paths: ["prompts/**", "config/models.yaml", "evals/**"]
jobs:
golden:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install -r requirements.txt
- run: pytest evals/ -q
env:
LLM_API_KEY: ${{ secrets.LLM_API_KEY }}
Use a dedicated API key with a spending cap for CI so a runaway job cannot surprise you at the end of the month. Also make sure pull requests from forks cannot read that secret.
Step 6: Compare against a baseline, not just a threshold
A pass or fail gate is a good start, but a baseline comparison tells you more. Save the per-case results from the main branch, then show on each pull request which cases changed from pass to fail and which went from fail to pass. Reviewers care about that diff far more than an overall percentage. A prompt change that fixes two cases and breaks one is a trade-off to discuss, and the report should make it visible.
Common mistakes
- Testing a copy of the prompt. If the test builds its own prompt, it can pass while production fails. Always call the same code path your app uses.
- Letting the dataset go stale. Add cases whenever you fix a bug, and prune ones that no longer reflect your product.
- Overfitting to the set. If you tune a prompt until every golden case passes, you may be memorizing them. Keep a small held-out group of cases that you look at only when deciding to ship.
- Checking only the happy path. Include refusals, malformed input, and cases where the right answer is "I don't know."
- Ignoring cost and latency. Record tokens and response time per case. A change that keeps quality but doubles the token count is still a regression for your budget.
A rollout plan you can finish this week
- Collect 30 cases: 10 from past failures, 10 from sampled traffic, 10 edge cases you write by hand.
- Add schema, string, and length checks for each one.
- Write the pytest runner against your real entry point and run it locally.
- Tag your 10 most important cases
coreand run only those on every pull request. - Add the path-filtered full run and a nightly job.
- From now on, every production bug adds a case before the fix merges.
None of this requires a special evaluation platform. A text file of cases, some plain assertions, and a CI job give you most of the safety, and you can add a dedicated tool later if the suite outgrows this setup.
Takeaway
Prompts, model versions, and retrieval settings are code in every way that matters, so they deserve tests. Start small, favor deterministic checks, run the expensive parts only when relevant files change, and turn every real failure into a permanent test case. Over a few weeks, the suite becomes the quickest way to change your LLM app with confidence.
Comments
Post a Comment