onetrace 0.2.0

Recipe for coding agents: instrument an existing Python pipeline with onetrace

You are adding onetrace to a codebase you did not write. Your job is to add recording, not to change behaviour. When you finish, the pipeline must produce exactly the same results as before, and every run must leave a record that onetrace-verify accepts.

Read all of this before editing. The rules in "Never" are hard rules. Fetch this file itself, as raw text, from https://oneproof.dev/onetrace/agents/instrument.md, and work from it, not from a summary of it.

The rules, in short (each is explained below):

  1. Add recording. Change nothing the pipeline computes: no signature, output or logic changes.
  2. Leave out trust=, output_trust=, rederivable= and note=, and ask the human about each.
  3. No secret, token or key in a decorator, config=, ot.constant or a decorated function's arguments: a step that takes a client holding a key gets a thin wrapper (§3).
  4. Never turn on signing, decorate library code, add network calls, anchoring or telemetry, or edit CI files unasked.
  5. List in @ot.run(stages=[...]) every stage the run's own module doesn't define when it is imported.
  6. A literal setting goes in config=; a setting from a flag, an argument or a variable goes in ot.constant.
  7. Use the real services. A stubbed run says so in its run id (ONETRACE_RUN_ID=stub-…), and isn't proof.
  8. Run every run in full, verify it with onetrace-verify --require-artifacts, and compare two runs.
  9. End with the report in §7, including your questions for the human.

0. Before you start

1. Find the run

A run is one call of one function that does the whole job once: answers one question, processes one document batch, handles one request. Find that function (often main, run, handle, answer, or the function the CLI or API route calls).

2. Find the stages

A stage is a function the run calls that does one distinct step whose result matters. Typical stages in AI pipelines:

Step Typical kind
load / read documents reader
clean, convert, normalise transformer
split into chunks transformer
embed embedder
build or update an index indexer
retrieve retriever
build the prompt formatter
call a model model-call
parse the model's output parser
act on it (write a file, call an API) actuator

Rules for choosing:

3. Add the decorators

import onetrace as ot

@ot.run(run_dir="runs/{run_id}")
def answer_question(question: str) -> str:
    docs = retrieve(question)
    return generate(question, docs)

@ot.stage("retrieve",
          instrument=ot.pkg("bm25-retriever", "rank_bm25", kind="retriever",
                            config={"top_k": 4}),
          files=["data/corpus.json"])
def retrieve(question: str) -> list[str]:
    ...

@ot.stage("generate",
          instrument=ot.pkg("model-call", "openai", kind="model-call",
                            config={"model": "gpt-4.1-mini"}))
def generate(question: str, docs: list[str]) -> str:
    ...

4. Leave the meaning to the human

These fields say what a stage means, and only a person may set them:

Do not set these. Leave them out. onetrace then records the most cautious value (externally-sourced; not re-derivable) and lists the field as undeclared on that stage, so everyone can see nobody stated it. At the end, list them as questions for the human (step 7), with what you observed, for example: "generate calls a hosted model: is it re-derivable? (usually no)".

5. Check behaviour is unchanged

  1. Run the project's test command. The same tests must pass as in step 0, with the same results. If anything differs, undo your change to that function and report it.
  2. Use the real services. If a service the pipeline needs is down (a vector database, a model API), stop and tell the human. Never substitute stubs or mocks silently. A run made with stubs must say so in its run id, and it doesn't count as proof. If the human agrees to stub a model API, do it on the command line, with no mock code inside the pipeline: start a local fake endpoint, point the client library at it with its base-URL environment variable, and give the run a stub id, in one line:
    ONETRACE_RUN_ID=stub-query-1 OPENAI_BASE_URL=http://127.0.0.1:8000/v1 python -m app "a question"
    
    OPENAI_BASE_URL is the openai package's variable; other client libraries name their own. Each run needs a new id: a run folder that already holds anything is refused.
  3. Run every run function in full (for a RAG pipeline, the whole ingest run and the whole query run, all stages), the way the project normally runs it or through the test that exercises it. A folder appears under runs/ for each, and when each run closes it prints that folder and the command that verifies it, on stderr. A shortened demo run that skips stages doesn't count.
  4. Verify it:
    onetrace-verify --require-artifacts runs/<run_id>
    
    Exit code 0 means the record is intact. The verifier's first line, on stderr, says so (record intact: …), or names the first failing row (record NOT intact: …), and under it a fix (…) line for each failing row says what to change and which section of the errors page explains it. Don't set ONETRACE_QUIET=1 while you check: it silences those lines. Anything else: read that row and its fix (…) line; the errors page explains each one. onetrace doctor checks the setup and says whether the most recent run verifies, not why a row failed. Fix the instrumentation, not the pipeline.
  5. Run it a second time with nothing changed, and compare: onetrace diff runs/<first> runs/<second>. Every stage should read same, except stages that call a hosted model, or read a service that changes by itself.
    • If any other stage differs, stop and find out why before anything else. The usual causes are ids generated fresh each run (uuid), timestamps written into outputs, absolute paths, and unordered sets or dict iteration.
    • Report it to the human with the stage and the field.
    • Until it's fixed, every comparison will blame that stage first and hide the real change.
  6. Only then make one deliberate change (for example chunk size) and compare again. The first difference must be the stage you changed. onetrace diff and onetrace localize name a setting recorded with ot.constant by its key ("constant chunk_size differs"); for a setting in config= they say the stage's config changed, not which key. To have the key named, record the setting with ot.constant.
  7. For a query pipeline, also compare two different questions. retrieve and the model's answer should normally differ. If they don't, report it: it usually means something was stubbed or cached. Also check that every stage that uses the question (retrieve, prompt building) lists it among its inputs.

6. Optional, only if the human asks

7. Report to the human

End with a short report:

Never