LLM adapters
An adapter runs an agent built in an LLM framework inside a simulated market on your machine. tradefloor sends the framework a JSON observation, the framework answers with a decision, and tradefloor checks the decision, places its orders in the simulation and scores the run. Each adapter is an ordinary agent with an act method, so it runs under tf.evaluate, tf.rank and World.
Three ways to connect a model, and which of them place orders
- LLM adapters, on this page, run the model inside a simulated market in your own Python process, and its decisions become orders in that simulation.
- The local MCP server gives an MCP client simulation and evaluation tools. No tool places an order.
- The hosted app, in beta, lets a model or a bot place orders in saved markets on app.tradefloor.dev over MCP, an HTTP API or an Alpaca-shaped API.
Setup
Each framework is an optional extra. tradefloor needs Python 3.11 or later.
| Framework | Install | Module | Minimum version |
|---|---|---|---|
| A plain Python function | pip install tradefloor | tradefloor.integrations.callable | none |
| OpenAI Agents SDK | pip install "tradefloor[openai-agents]" | tradefloor.integrations.openai_agents | openai-agents 0.22 |
| PydanticAI | pip install "tradefloor[pydantic-ai]" | tradefloor.integrations.pydantic_ai | pydantic-ai-slim 2.36 |
| LangGraph | pip install "tradefloor[langgraph]" | tradefloor.integrations.langgraph | langgraph 1.2 |
| FinRobot | pip install "tradefloor[finrobot]" | tradefloor.integrations.finrobot | Python 3.11 exactly |
A replay of a recorded run needs none of the extras.
A first run with no API key
The callable adapter wraps any function that takes the observation payload and returns a decision. A function that answers without a model is the fastest way to check the loop before any key or extra is involved.
import json
import tradefloor as tf
from tradefloor.integrations.callable import callable_agent
def ask_model(payload):
# stands in for a model call: the payload is what a model is shown
first = payload["assets"][0]
order = {"symbol": first["symbol"], "side": "BUY", "quantity": 100}
return json.dumps({"actions": [order], "rationale": "add 100 a day"})
adapter = callable_agent(ask_model)
market = tf.Universe.random(12, seed=4242)
card = tf.evaluate({"model": adapter}, seed=4242, universe=market, days=5)["model"]
print(card)
print(len(adapter.record), "decisions")Scorecard('model', pnl=1,027, return=+0.10%, trades=5, impact=+0.09bps, sharpe=n/a (short run), vol=0.4%, in_market=100%, exposure=0.02x)
5 decisionsReplace the body of ask_model with a call to your model, and parse its text into a decision with postprocess=, as in Recording and replay. The record in adapter.record holds one entry per decision: the payload, the exact input, the raw response, the validated decision and the orders.
A first run in each framework
Each block below runs on its own with no API key, because it hands the adapter an offline model the framework ships for testing. They need the extra, and were tested with openai-agents 0.23.1, pydantic-ai-slim 2.53.0 and langgraph 1.2.12 on tradefloor 0.8.6.
PydanticAI
TestModel answers every request with the arguments it is given. Pass your own model through the adapter's model= argument the same way.
import tradefloor as tf
from pydantic_ai import Agent
from pydantic_ai.models.test import TestModel
from tradefloor.integrations.pydantic_ai import PydanticAIAdapter
# a scripted model: no API key, the same answer at every decision
offline = TestModel(custom_output_args={
"actions": [{"symbol": "AAA", "side": "BUY", "quantity": 200}],
"rationale": "add to AAA"})
pm = Agent(instructions="You manage a small portfolio.")
adapter = PydanticAIAdapter(pm, model=offline)
market = tf.Universe.random(12, seed=4242)
card = tf.evaluate({"pm": adapter}, seed=4242, universe=market, days=5)["pm"]
print(card)
print(len(adapter.record), "decisions")Scorecard('pm', pnl=2,055, return=+0.21%, trades=5, impact=+0.18bps, sharpe=n/a (short run), vol=0.7%, in_market=100%, exposure=0.03x)
5 decisionsFor a live run, build the Agent with your provider's model, as PydanticAI's documentation describes, and leave out model=. Your agent's tools, deps and instructions are passed through unchanged, and the adapter binds the decision schema as the output type for each run.
OpenAI Agents SDK
ScriptedModel from agents.testing answers each call with a function of the input. payload_of reads back the observation the adapter sent.
import json
import tradefloor as tf
from agents import Agent
from agents.testing import ModelStep, ScriptedModel, assistant_message
from tradefloor.integrations.openai_agents import OpenAIAgentsAdapter, payload_of
def answer(call):
# reads the same payload a real model is sent
first = payload_of(call)["assets"][0]
order = {"symbol": first["symbol"], "side": "BUY", "quantity": 200}
return [assistant_message(json.dumps({"actions": [order],
"rationale": "add"}))]
offline = ScriptedModel([ModelStep.respond(answer)] * 5)
pm = Agent(name="Portfolio Manager", instructions="You manage a small portfolio.")
adapter = OpenAIAgentsAdapter(pm, mode="live", model=offline)
market = tf.Universe.random(12, seed=4242)
card = tf.evaluate({"pm": adapter}, seed=4242, universe=market, days=5)["pm"]
print(card)
print(len(adapter.record), "decisions")Scorecard('pm', pnl=2,055, return=+0.21%, trades=5, impact=+0.18bps, sharpe=n/a (short run), vol=0.7%, in_market=100%, exposure=0.03x)
5 decisionsOpenAIAgentsAdapter defaults to mode="replay", so a live run names mode="live", and openai_agent(agent) is the same adapter with live as its default. For a live run, give the Agent a model your account can use, set OPENAI_API_KEY and leave out model=. The adapter turns the SDK's tracing off for each run it starts unless you pass tracing=True, because tracing is on by default in the SDK and sends traces to OpenAI.
LangGraph
The adapter takes any compiled graph. By default it sends the payload under observation and as messages, and reads the decision from a decision key, an actions key or the last message.
from typing import Any, TypedDict
import tradefloor as tf
from langgraph.graph import END, START, StateGraph
from tradefloor.integrations.langgraph import LangGraphAdapter
class State(TypedDict, total=False):
observation: dict[str, Any]
decision: dict[str, Any]
def decide(state: State) -> State:
# a rule in place of a model: buy 200 of the first name when flat
first = state["observation"]["assets"][0]
if first["position"] > 0:
return {"decision": {"actions": [], "rationale": "holding"}}
order = {"symbol": first["symbol"], "side": "BUY", "quantity": 200}
return {"decision": {"actions": [order], "rationale": "open"}}
builder = StateGraph(State)
builder.add_node("decide", decide)
builder.add_edge(START, "decide")
builder.add_edge("decide", END)
adapter = LangGraphAdapter(builder.compile())
market = tf.Universe.random(12, seed=4242)
card = tf.evaluate({"graph": adapter}, seed=4242, universe=market, days=5)["graph"]
print(card)
print(len(adapter.record), "decisions")Scorecard('graph', pnl=544, return=+0.05%, trades=1, impact=+0.04bps, sharpe=n/a (short run), vol=0.1%, in_market=100%, exposure=0.01x)
5 decisionsA node that calls a model makes the graph live, and nothing else changes. The graph decides five times and trades once, because after the first day it holds the name and answers with an empty action list, which is a decision to change nothing.
FinRobot
The FinRobot adapter predates the shared layer and is documented in the reference, FinRobot adapter. It needs the finrobot extra on Python 3.11.
Decisions and model calls
An adapter asks its framework once every every decision steps, and the default is 6. obs.step counts steps across the whole run, so with the default 6 steps a day that is one decision a simulated day, at each day's first step. With steps_per_day=3 the same default asks every second day. On the steps between decisions the adapter records the prices it saw and returns no orders.
One decision can cost several calls to the model provider, so a 20-day run is 20 decisions and can be many more provider calls:
- A framework turn is one model call, and an agent that calls tools takes several turns per decision. The OpenAI Agents adapter caps a decision at
max_turns=6, and the PydanticAI adapter atrequest_limit=8requests. - PydanticAI retries a schema violation inside its own loop, and each retry is another request. The OpenAI Agents SDK 0.22 makes no client-side retry on a malformed answer.
- A LangGraph graph makes whatever calls its nodes make.
The observation payload
The framework is shown the serialized payload and never the observation object. The payload is an allowlist written out field by field:
| Key | What it holds |
|---|---|
step, day, steps_per_day | The clock. step counts the whole run. |
macro | The published figures: the VIX, the policy rate, the corporate bond yield, inflation and the business-cycle phase as announced, late. |
assets | Per name: the symbol, price, one-day and five-day returns and volatility from the prices the adapter has seen, the best bid and ask, average daily volume, max_order_shares (the participation cap), the position held, and any fundamentals you pass with fundamentals=. |
portfolio | Cash, net worth, leverage, max_leverage, buying_power and the waiting limit orders. |
It holds no fair value, no factor attribution and no macro path the run has not reached. That holds at every access level: the adapter's own act receives the read-only market view by default, or the live engine under trusted_agents=True, as What the agent sees sets out, and in both cases the framework receives the payload alone.
The decision is a list of actions, each with a symbol, a side (BUY, SELL, HOLD or CANCEL), a positive quantity in shares, and an optional limit_price that makes it a limit order. A bad action is refused on its own and listed in the scorecard's errors. An order over max_order_shares, 2% of the name's average daily volume by default, is cut to the cap, and an order below one share is dropped. Bringing an LLM agent has the full contract.
Reference
LLM adapters and MCP has every adapter's signature, the shared layer, error types, replay rules and the framework version each adapter was written against. When a run fails or makes no trades, Troubleshooting lists the messages and fixes.