Skip to the page
GUIDES/RECORD AND REPLAY

Record and replay

A seed replays a tradefloor market exactly. A model behind an API answers differently every time and costs money on every call. Recording what the model answered, and replaying the run from that recording, makes a result from an LLM agent reproducible and lets a reader check it without an API key. The adapters' reference is LLM adapters and MCP, and setting an adapter up is covered in LLM adapters.

Every code block on this page runs on its own except the two lines in Saving a recording, which continue the recording built in Recording and replay. None of them calls a model, because a plain function stands in for one, and none needs an optional extra.

Recording and replay

An adapter in mode="live" with a recorder calls the model and writes each answer to the recorder, a transcript. The same adapter in mode="replay" with that recording as its transcript answers from it and never calls the model.

import json

import tradefloor as tf
from tradefloor.integrations.callable import callable_agent
from tradefloor.integrations.common import AdapterInfo, Transcript, digest

PROMPT = "You manage a portfolio. Answer with one JSON decision."

def ask_model(payload):
    # your model call goes here; this stand-in answers in text
    first = payload["assets"][0]
    order = {"symbol": first["symbol"], "side": "BUY", "quantity": 400,
             "limit_price": first["best_bid"]}
    return json.dumps({"actions": [order], "rationale": "bid at the touch"})

def to_decision(raw, payload):
    decision = json.loads(raw)
    for action in decision["actions"]:
        action["quantity"] = min(action["quantity"], 100)   # your risk cap
    return decision

info = AdapterInfo(framework="callable", instructions_digest=digest(PROMPT))
market = tf.Universe.random(12, seed=4242)

recording = Transcript()
live = callable_agent(ask_model, postprocess=to_decision, info=info,
                      mode="live", recorder=recording)
first = tf.evaluate({"m": live}, seed=4242, universe=market, days=5)["m"]

replay = callable_agent(postprocess=to_decision, info=info, mode="replay",
                        transcript=recording)
again = tf.evaluate({"m": replay}, seed=4242, universe=market, days=5)["m"]
print("replay matches", first.pnl == again.pnl)
print(len(recording), "recorded answers")
print({key: recording.meta[key] for key in
       ("model_preset", "observation_schema_version", "decision_schema_version")})
replay matches True
5 recorded answers
{'model_preset': 'pt-v20', 'observation_schema_version': '1', 'decision_schema_version': '2'}

An adapter asks its model every 6 decision steps by default, once a simulated day, so five days are five decisions and five recorded answers. ask_model returns the model's raw text and to_decision turns it into a decision. The transcript records the raw text, so the parsing and the risk cap in to_decision run again on every replay. Code inside ask_model after the model call is recorded as its output and does not run on replay.

Replay checks

Each recorded answer is keyed by a digest of the exact input the model was sent, never by a step number. A replay that reaches an input with no recorded answer stops with ReplayMiss and names the step, so a changed experiment cannot quietly reuse the answers given to an old one.

ChangeWhat happens on replay
Seed, universe or cadenceThe input differs, so the first lookup misses and the run stops with ReplayMiss.
PresetRefused with ReplayMiss naming both presets before any lookup. Recording and replay must run on the same preset, set with model= on evaluate or World.
Observation or decision schema versionRefused before the first lookup, naming both versions.
The model's instructionsRefused with ValidationError when the adapter is built, if the recording carries an instructions digest. The callable adapter carries one only when its AdapterInfo does, as here.

Instructions get their own check because most adapters send them outside the input the replay is keyed on, so a changed prompt would otherwise match every key. The LangGraph adapter renders its instructions into the input, so there a changed prompt is a lookup miss.

import json

import tradefloor as tf
from tradefloor.integrations.callable import callable_agent
from tradefloor.integrations.common import (AdapterInfo, ReplayMiss,
                                            Transcript, digest)

PROMPT = "You manage a portfolio. Answer with one JSON decision."

def ask_model(payload):
    first = payload["assets"][0]
    order = {"symbol": first["symbol"], "side": "BUY", "quantity": 100}
    return json.dumps({"actions": [order], "rationale": "add"})

info = AdapterInfo(framework="callable", instructions_digest=digest(PROMPT))
market = tf.Universe.random(12, seed=4242)
recording = Transcript()
live = callable_agent(ask_model, info=info, mode="live", recorder=recording)
tf.evaluate({"m": live}, seed=4242, universe=market, days=5)

# the same recording, replayed on another seed
replay = callable_agent(info=info, mode="replay", transcript=recording)
try:
    tf.evaluate({"m": replay}, seed=4243, universe=market, days=5)
except ReplayMiss:
    print("seed 4243: ReplayMiss")

# the same recording, under an edited prompt
edited = AdapterInfo(framework="callable",
                     instructions_digest=digest(PROMPT + " Be bold."))
try:
    callable_agent(info=edited, mode="replay", transcript=recording)
except tf.ValidationError:
    print("edited prompt: ValidationError")
seed 4243: ReplayMiss
edited prompt: ValidationError

evaluate lets ReplayMiss end the run rather than scoring it as an error, because a broken recording scored as an error would read as an agent that held cash from that step on. Replay mismatches gives the fix for each failure.

Saving a recording

A transcript is plain JSON. recording.save(path) writes it and Transcript.load(path) reads it back, and to_json() and Transcript.from_json(text) do the same with a string:

recording.save("run-4242.json")
recording = Transcript.load("run-4242.json")

A recording made before 0.8.5 has no schema versions, and its first lookup misses on 0.8.5 or later because the payload it was keyed on changed. Replay it on the release that made it.

Recording the market

The transcript reproduces the model. The market is reproduced from its seed and its order log, and a manifest, RunManifest, carries both, with the package version, the preset and a digest the replay has to match. evaluate and rank write no manifest, so run the agent in a World to get one.

import json

import tradefloor as tf
from tradefloor.integrations.callable import callable_agent
from tradefloor.integrations.common import AdapterInfo, Transcript, digest

PROMPT = "You manage a portfolio. Answer with one JSON decision."

def ask_model(payload):
    first = payload["assets"][0]
    order = {"symbol": first["symbol"], "side": "BUY", "quantity": 100}
    return json.dumps({"actions": [order], "rationale": "add"})

info = AdapterInfo(framework="callable", instructions_digest=digest(PROMPT))
market = tf.Universe.random(12, seed=4242)
recording = Transcript()
agent = callable_agent(ask_model, info=info, mode="live", recorder=recording)

world = tf.World(seed=4242, universe=market, agent=agent)
world.run(5)
manifest = world.manifest()
text = manifest.to_json()                        # publish this beside the transcript

engine = tf.RunManifest.from_json(text).reproduce()
print("market reproduced", engine.state_hash() == world.engine.state_hash())
print(manifest.result["days"], "days,", len(recording), "recorded answers")
market reproduced True
5 days, 5 recorded answers

reproduce() rebuilds the engine from the manifest and raises ValidationError naming the component that disagreed. It checks the market and carries no score. RunManifest lists what a manifest holds.

Publishing a result

To let a reader check a result from an LLM agent, publish:

  • the transcript, which replays the model's answers
  • the manifest, which replays the market
  • adapter.provenance(), which names the framework, the model, the instructions digest and the settings that change what a decision can be
  • the agent's code and the call that ran it, with its seeds

To show a score was not tuned to its seeds, publish tf.commit(seeds, salt) before the run, then the seeds and the salt after it. tf.reveal checks the three agree. Bringing an LLM agent covers the fixed seven-market battery that tf.fingerprint.fingerprint runs to tell two versions of an agent apart.

Next steps

  • LLM adapters sets up the OpenAI Agents SDK, PydanticAI, LangGraph and FinRobot adapters, which record and replay the same way.
  • Replay and recording in the reference covers each adapter's instruction check and what a record entry holds.