Skip to the page
API REFERENCE/LLM ADAPTERS AND MCP

LLM adapters and MCP

Reference for the LLM adapters and the local MCP server. The guides set each one up and run it: LLM adapters, Local MCP server and Record and replay. A summary of each comes first, then the shared layer and the four framework adapters: signatures, input and output mapping, errors, replay, tracing and the framework version each was written against.

Bringing an LLM agent

An adapter runs an agent built in another framework inside a tradefloor market. The framework reads a JSON observation and answers with a decision, and tradefloor checks the decision, sends the orders to the book and scores the result. Each adapter is an ordinary agent with an act method, so it runs under tf.evaluate, tf.rank and World, and two frameworks run on the same seed meet the same market. LLM adapters sets each framework up and runs it with no API key, and Record and replay records a run and replays it without the model.

FrameworkModuleInstall extra
Plain Python functiontradefloor.integrations.callablenone
OpenAI Agents SDKtradefloor.integrations.openai_agentsopenai-agents
PydanticAItradefloor.integrations.pydantic_aipydantic-ai
LangGraphtradefloor.integrations.langgraphlanggraph
FinRobottradefloor.integrations.finrobotfinrobot, Python 3.11 only

An adapter asks its framework every every decision steps, 6 by default, which at the default 6 steps a day is one decision a simulated day. One decision can take several calls to the model provider, through tool calls, framework turns and retries, as Decisions and model calls sets out.

The decision is the same for every adapter: a list of actions, each with a symbol, a side (BUY, SELL, HOLD or CANCEL), a positive quantity, and an optional limit_price that makes it a limit order. A bad action, such as an unknown symbol or a negative quantity, is refused on its own and recorded in the scorecard's errors, and the rest trade. An order over the participation cap, 2% of the name's average daily volume by default, is cut to the cap. The observation payload never carries fair value, the attribution of a move or the macro path the run has not reached.

A seed replays the market exactly, and a Transcript replays the model. Each entry is keyed by the payload the model was shown, so a changed roster, seed or cadence stops a replay with ReplayMiss, and a replay under different instructions is refused before the market opens. To publish a result from an external agent, record both sides: adapter.provenance() for the framework, model and settings, and a RunManifest for the market.

tf.fingerprint.fingerprint(agent) runs an agent on a fixed battery of seven 120-day markets, one per shipped scenario, and hashes what it ordered, so two versions of an agent can be checked for whether they behave the same. tf.battery() returns that battery and BATTERY_VERSION names it. To show a score was not tuned to its seeds, publish tf.commit(seeds, salt) before the run, then the seeds and the salt, which tf.reveal checks. tf.sealed_battery(seeds, salt) builds the battery on them.

The MCP server

tradefloor-mcp offers the simulator as tools a model calls over the Model Context Protocol, on your machine over stdio. Local MCP server installs it, registers it with a client and makes a first call.

Over MCP a model studies the market from outside. No tool places an order, and a tool that runs a market gives the same result for the same arguments. A model that trades inside the market is an agent, run through an adapter, and the hosted app serves its own MCP tools, which do place orders.

ToolWhat it does
describe_simulatorwhat the simulator is, what it is measured to reproduce and what it cannot do
check_envelopewhether a question is inside the validated scope, before anything runs
validate_strategychecks a strategy spec and returns its fingerprint
build_universebuilds a roster: generated, concentrated in some sectors, or written by hand
build_scenariocomposes a scenario and shows what it resolves to
list_scenariosthe shipped scenarios, the constructors and the targets
evaluate_strategiesruns strategies and the five baselines on one market
rank_strategiesruns them across seeds, with a paired sign test between each pair
run_stress_scenariowhat a scenario does to each strategy, against the same market without it
explain_price_movea price move split into the eleven factors that sum to it
explainthe random draws behind one day for one company
start_job, check_jobslow work in the background, up to 252 days

A strategy is a JSON StrategySpec, never code, and every result carries a caveats list and a provenance block with the version, preset, seed and roster, so a model summarizing it sees the limits too. A direct call runs at most 60 days and a background job 252, the horizon the model is validated over. No tool takes a preset: every run uses pt-v20 and names it. Limits lists the rest of the caps.

The shared layer

tradefloor.integrations.common holds the parts every adapter needs that belong to no single framework. These are the observation allowlist, the decision schema and the model built from it, two-stage validation, transcripts and replay, and the base class an adapter completes.

NameWhat it is
serialize_observationThe observation allowlist, the fields a framework is shown. The framework gets this payload and never the Observation itself. By default the Observation's .engine is the read-only market view, which serves none of the hidden state. Under trusted_agents=True it is the live engine, and through it fair value, the eleven-way attribution of every move, each company's mispricing and the macro path the run has not reached yet. The payload is the same allowlist either way.
decision_schemaThe one definition of a valid decision. decision_model builds the Pydantic model from it.
parse_decisionThe first validation stage. It checks the answer against the schema. An unknown top-level key, or no actions list, refuses the whole decision with DecisionError. A bad action, such as one with an unknown field, a negative quantity or a limit order with no price, is refused on its own: it goes into the returned Decision's refused list with its reason, and the other actions stand.
orders_fromThe second validation stage. It checks each symbol against the listed universe and each size against the participation cap, the largest fraction of average daily volume one order may take, and turns the actions into orders: a share count for a market order, a tf.Limit for an action with a limit_price, and a tf.Cancel for CANCEL. An order over the cap is clipped, and the clip is recorded. An unlisted symbol goes into the refused list it is given, and raises MarketRefusalError when it is given none.
TranscriptRecord and replay. Each exchange is keyed by a digest, a hash of the exact input sent to the framework, and meta names the market the recording was made in.
run_syncThe one supported bridge from sync code to async code. It runs a coroutine on another thread, so context variables the caller set do not carry into it.
FrameworkAdapterThe base class an adapter completes. ReplayMixin adds the record-and-replay path. Every adapter shares the constructor arguments info, every, fundamentals, max_participation and arm.

The module exports seven constants, OBSERVABLE_MACRO, SIDES, ORDER_TYPES, HISTORY_STEPS, MAX_PARTICIPATION, DECISION_SCHEMA_VERSION ("2" since 0.8.5) and OBSERVATION_SCHEMA_VERSION ("1"). Both versions are frozen for the 0.8.x line. The observation payload lists the payload, and Bringing an LLM agent lists the decision. It also exports three classes, Action, Decision and AdapterInfo, and three functions, require, digest and replay_response. It also exports the preset check the built-in adapters run, so an adapter you write yourself can run it too. The check is preset_of(obs), stamp_preset(recorder, obs), refuse_a_changed_preset(transcript, preset), stamp_artefact(meta) and the constant PRESET_VECTOR_KEY.

Error behavior

Every error below is a subclass of tradefloor.ValidationError. A bad answer from the agent raises DecisionError, and a failure in the framework raises FrameworkError. The two are kept apart so that a weak agent is not mistaken for an unreliable network.

NameMeaning
IntegrationErrorThe root of this error family.
MissingDependencyErrorAlso an ImportError. require() raises it.
FrameworkErrorThe framework call itself failed, for example with a timeout, a transport failure or a budget stop.
DecisionErrorThe agent's output could not be read as a decision. The error names the step and the day.
MarketRefusalErrorA DecisionError. The decision is well formed, and this market cannot take it. The adapters pass a refused list, so one bad action is refused on its own and the rest of the decision trades.
ReplayMissA DecisionError. The recording has no answer for this input, or it was made in a different market. World re-raises it under on_refusal="skip", so a replay that cannot answer stops the run instead of counting against the agent.

No adapter turns a failure into an empty decision. An empty decision scores as trades=0 with an empty error column, which is also how an agent that chose not to trade scores. If failures became empty decisions, an agent that failed and an agent that declined would score the same.

Replay and recording

A Transcript keys each exchange by a digest of the exact input sent to the framework, never by a step number. If you change the roster, the seed or the cadence, the key has no recorded answer, so the run stops with ReplayMiss and names the step. With a step-number key, a changed experiment would get the answers recorded for the old one, and nothing in the output would show it.

Instructions are checked separately, because most adapters do not send them in the keyed input. LangGraphAdapter renders its instructions into that input, so a change stops the run as above. The OpenAI Agents, PydanticAI and FinRobot adapters write a digest of their instructions to meta["instructions_digest"], and a replay under different instructions raises ValidationError when the adapter is built. CallableAgentAdapter cannot see the prompt inside your function, so it checks only when its AdapterInfo carries an instructions_digest.

adapter.record holds one entry per decision, with the digest, the exact input, the raw response, the validated decision, the orders, any participation clips and any refused actions with their reasons. A limit order is held as {"quantity", "limit_price"}. So the whole path from observation to order is in one place, in a replayed run as in a live one.

Recorded preset

Every price in an observation comes from the preset, the named set of model coefficients the engine runs. So a recording only replays against the preset it was made in, and the recording names that preset. On the first recorded exchange, every adapter writes the running engine's fingerprint to meta["model_preset"] and the full parameter vector to meta["model_preset_vector"]. The fingerprint is a preset name such as pt-v19 or custom-XXXXXXXX. Transcript.save writes recorded_utc. For a transcript that never ran against an engine, it also writes the shipped default as model_preset. finrobot.Transcript.save writes the same two fields.

A replay also refuses a different market. replay_response(transcript, key, *, step, day, preset=None) compares the recorded preset with preset before it looks up the digest, and every adapter passes the running engine's preset. A different name raises ReplayMiss, naming both. To fix it, replay on the recorded preset with World(..., model=...) or evaluate(..., model=...), or record the run again. The same preset name can also carry different values, which happens when a preset is cut again between builds. That also raises ReplayMiss, and the error lists the dials (model parameters) that moved.

A recording made before 0.8.0 has no model_preset, and a replay of it warns and goes on. A recording's meta also carries observation_schema_version and decision_schema_version, and a replay of a recording made under another payload version stops before the first lookup and names both versions. A recording made before 0.8.5 carries neither, and its first lookup misses, because the payload it was keyed on changed. Replay it on the release that made it.

Generic callable

CallableAgentAdapter wraps a plain Python function so that evaluate can run it, and callable_agent(fn, **kwargs) is the convenience constructor. An async function runs through common.run_sync. The function gets the serialized payload, never the Observation.

CallableAgentAdapter(fn=None, *, name, info, every=6, mode="live", transcript, recorder, prior, postprocess, ...). fn may be left out in replay mode, which never calls it.

postprocess(raw, payload) runs on whatever fn returned, in a live run and in a replay, and its return is the decision. The transcript records fn's return, the raw model response, so parsing, risk checks and sizing written in postprocess are exercised by every replay. Code inside fn after the model call is recorded as its output and never runs on replay. Without postprocess, fn's return is the decision.

The replay key is the payload alone. Pass info=AdapterInfo(framework="callable", instructions_digest=digest(PROMPT)) when you record and when you replay. The digest is written to the transcript's meta, and a replay built with a different digest raises ValidationError before the market opens. With no AdapterInfo, a replay under an edited prompt runs to the end on the recorded answers. Replay checks has a worked example.

OpenAI Agents SDK adapter

OpenAIAgentsAdapter(agent, *, mode="replay", transcript, recorder, model, brief, max_turns=6, tracing=False, run_id, ...). The convenience constructor openai_agent(agent, **kwargs) defaults to mode="live". The module also exports payload_of(call), BRIEF, DISTRIBUTION and EXTRA. The adapter adds two methods to the base, ask(obs, payload) and input_items(payload). state() adds max_turns and brief_digest, because both change what a decision can be.

Compatibility

The adapter binds the decision schema as the output type on a copy made with Agent.clone(...). Your agent is left unchanged, and its instructions, tools, model settings, hooks, handoffs and guardrails all carry over. An agent that already declares its own output_type is refused, and the message says how to opt in. A run that ends on a different agent through a handoff is refused by name, because that agent has an output type of its own.

Runtime behavior

The adapter calls Runner.run through the shared async bridge. On 0.22.0, Runner.run_sync cannot run inside an existing event loop, so it does not work from a notebook. max_turns caps one decision at six model calls. The SDK's default of ten suits interactive use, and a loop that runs at every cadence step of every experiment arm needs the lower cap. The adapter imports the SDK only inside the method that calls it, so a replayed run needs neither the package nor the time its import takes.

Retry and validation behavior

On 0.22.0, the SDK makes one model call on a malformed answer and raises ModelBehaviorError. It does not retry on the client side. Binding the decision model puts the side enum, the non-negative quantity and additionalProperties: false into the schema the provider sees, which makes an invalid decision less likely. The binding adds no repair loop, so write your error handling for this adapter as if there were no retries.

Error behavior

The adapter treats seven SDK exceptions as the agent's own outcome and turns each into a DecisionError: ModelBehaviorError, ModelRefusalError, MaxTurnsExceeded and all four guardrail tripwires. By default the SDK removes model text from its own messages, so its message alone cannot say which decision point failed. A tripped guardrail is not converted to a hold. Every other exception stays a FrameworkError with the exception chain intact.

Tracing

SDK tracing is on by default and sends traces to OpenAI. This adapter passes tracing_disabled=True on every run it starts, unless you gave tracing=True. It does this per run instead of through the SDK's process-wide switch, so tracing for other code in the process is left alone.

Version notes

The minimum version is openai-agents>=0.22, the version the adapter was written against. A minimum at the major version would allow releases the adapter has not been tested against. The package imports as agents and supports Python 3.10 through 3.14. tradefloor needs 3.11 or later, so every Python that runs tradefloor runs the SDK.

PydanticAI adapter

PydanticAIAdapter(agent, *, deps, mode="live", model, transcript, recorder, instructions, bind_output_type=True, request_limit=8, ...). The module also exports UsageLimitReached, render(payload), MANDATE and MANDATE_VERSION.

Compatibility

You pass in a built agent, and the adapter does not modify it. Its deps_type, tools, RunContext usage, instructions, toolsets and output type keep working, and deps reaches run(deps=...) unchanged. PydanticAI has one dependency slot. Every tool reads it through RunContext.deps, and nothing checks its type at runtime, so an adapter that put its own payload there would hand your tools an object of the wrong type. The adapter puts nothing in it, so a tool cannot query the observation from inside a decision. A tool that needs the day's prices reads them from a holder on your own deps object.

Decision mapping

The adapter binds the shared decision model as the run's output type, so the side enum, the non-negative share count and the required actions list are in the schema the model sees. The binding applies to that run only, and your agent's own output type is untouched. PydanticAI does not allow an override of the output type on an agent with an @agent.output_validator. In that case the adapter says so and points to bind_output_type=False. With that set, your output type stays and tradefloor still validates what it produces.

Retry and validation behavior

PydanticAI's own retry loop catches a schema violation and corrects it within the turn. The adapter files UnexpectedModelBehavior as a DecisionError, because an agent that used up its retries without a valid decision answered badly. A run that hits its request limit raises UsageLimitReached, a FrameworkError subclass, so you can catch a deliberate budget stop by name.

Runtime behavior

Pass an offline model through the adapter's model= argument instead of Agent.override. Agent.override is built on context variables, and those do not carry across the shared async bridge. A test suite can set models.ALLOW_MODEL_REQUESTS = False. A blocked request then raises a plain RuntimeError, so a test that expects a framework exception will miss it.

Observability

PydanticAI records no traces by default, and the adapter turns nothing on. In the supported version, you set up tracing through PydanticAI's own current APIs, either logfire.configure() with logfire.instrument_pydantic_ai(), or Agent.instrument_all(). 2.36.0 has no instrument= constructor argument.

Version notes

The minimum version is pydantic-ai-slim>=2.36. The slim package has the pydantic_ai module without the provider SDKs that the full package adds, and no adapter imports any of those. TestModel and FunctionModel are both in slim.

LangGraph adapter

LangGraphAdapter(runnable, *, mode="live", transcript, recorder, input_builder, output_parser, instructions, config, thread_id, ...), with langgraph_agent(runnable, **kwargs) as the convenience constructor. The module also exports default_input_builder, default_output_parser, render, GraphInterruptedError, INSTRUCTIONS, DEFAULT_INPUT_KEYS and INTERRUPT_KEY.

Compatibility

The adapter accepts any object with an invoke method and does not check its class. Runnable is a nominal ABC with no __subclasshook__, so an isinstance check would reject a plain object with a working invoke, and a deterministic test double is that kind of object. An object with only ainvoke runs through common.run_sync, and an uncompiled StateGraph is refused by name.

Input mapping

The default input carries both input shapes at once, observation for a structured graph and messages for the MessagesState shape. A key the graph's state schema does not declare is dropped before any node runs, so one default works for both. A TypedDict state gets no schema check at the graph boundary, so a mismatch shows up as an IndexError or a bare KeyError inside your own node. Use input_builder when the graph expects a different state schema.

Output mapping

A graph returns its whole state, so the adapter has to take the decision out of it. The default parser reads a Decision, an interrupted state, a state carrying actions, a state carrying decision, or the MessagesState shape. It refuses anything else by name.

Interrupts

An interrupt raises GraphInterruptedError, a DecisionError subclass, which names the question that went unanswered. An interrupt means pause now and resume later. A market loop has nowhere to resume into, because the order book, the macro path and the variance process move on as soon as act returns. On 1.2.11, a GraphInterrupt never escapes invoke and arrives inside the state as __interrupt__. The parser checks for it before it looks for decision, because a checkpointed thread can carry a decision written on an earlier step. Run a human-in-the-loop graph separately, and give tradefloor a graph that decides.

Observability

The adapter puts the run's identity in config["metadata"] and adds a "tradefloor" tag, merged into any RunnableConfig you pass. It uses metadata because a node receives tags, metadata and recursion_limit from the invoking config, while the tracer consumes run_name. LangSmith tracing stays off unless one of its environment variables is true. Every exported run carries tradefloor_run_id. The adapter also stamps tradefloor_arm, tradefloor_day, tradefloor_step and tradefloor_decision_schema on each decision. Turning tracing on sends the rendered observation to LangSmith.

Version notes

The minimum version is langgraph>=1.2. That one requirement is enough, because langchain-core is a hard dependency of it. LangGraph needs Python 3.10 or later, so with tradefloor it runs on 3.11 and later. create_react_agent is deprecated in LangGraph 1.x and points at a package this extra does not install. A plain StateGraph has no deprecated import, and the adapter treats both the same way.

FinRobot adapter

The FinRobot integration is older than the shared layer, which was built from it. Its DecisionError inherits from common.DecisionError instead of directly from ValidationError. It is still importable, it is still a ValidationError, and it can also be caught as the shared error. It also runs the shared layer's preset check. A FinRobot recording names its market in meta, and a replay against a different one raises ReplayMiss.

The finrobot extra requires exactly Python 3.11, because FinRobot declares >=3.10, <3.12 and tradefloor needs >=3.11. Replaying a recorded run needs none of it.