Skip to the page
GUIDES/LLM ADAPTERS

LLM adapters

An adapter runs an agent built in an LLM framework inside a simulated market on your machine. tradefloor sends the framework a JSON observation, the framework answers with a decision, and tradefloor checks the decision, places its orders in the simulation and scores the run. Each adapter is an ordinary agent with an act method, so it runs under tf.evaluate, tf.rank and World.

Three ways to connect a model, and which of them place orders

  • LLM adapters, on this page, run the model inside a simulated market in your own Python process, and its decisions become orders in that simulation.
  • The local MCP server gives an MCP client simulation and evaluation tools. No tool places an order.
  • The hosted app, in beta, lets a model or a bot place orders in saved markets on app.tradefloor.dev over MCP, an HTTP API or an Alpaca-shaped API.

Setup

Each framework is an optional extra. tradefloor needs Python 3.11 or later.

FrameworkInstallModuleMinimum version
A plain Python functionpip install tradefloortradefloor.integrations.callablenone
OpenAI Agents SDKpip install "tradefloor[openai-agents]"tradefloor.integrations.openai_agentsopenai-agents 0.22
PydanticAIpip install "tradefloor[pydantic-ai]"tradefloor.integrations.pydantic_aipydantic-ai-slim 2.36
LangGraphpip install "tradefloor[langgraph]"tradefloor.integrations.langgraphlanggraph 1.2
FinRobotpip install "tradefloor[finrobot]"tradefloor.integrations.finrobotPython 3.11 exactly

A replay of a recorded run needs none of the extras.

A first run with no API key

The callable adapter wraps any function that takes the observation payload and returns a decision. A function that answers without a model is the fastest way to check the loop before any key or extra is involved.

import json

import tradefloor as tf
from tradefloor.integrations.callable import callable_agent

def ask_model(payload):
    # stands in for a model call: the payload is what a model is shown
    first = payload["assets"][0]
    order = {"symbol": first["symbol"], "side": "BUY", "quantity": 100}
    return json.dumps({"actions": [order], "rationale": "add 100 a day"})

adapter = callable_agent(ask_model)
market = tf.Universe.random(12, seed=4242)
card = tf.evaluate({"model": adapter}, seed=4242, universe=market, days=5)["model"]
print(card)
print(len(adapter.record), "decisions")
Scorecard('model', pnl=1,027, return=+0.10%, trades=5, impact=+0.09bps, sharpe=n/a (short run), vol=0.4%, in_market=100%, exposure=0.02x)
5 decisions

Replace the body of ask_model with a call to your model, and parse its text into a decision with postprocess=, as in Recording and replay. The record in adapter.record holds one entry per decision: the payload, the exact input, the raw response, the validated decision and the orders.

A first run in each framework

Each block below runs on its own with no API key, because it hands the adapter an offline model the framework ships for testing. They need the extra, and were tested with openai-agents 0.23.1, pydantic-ai-slim 2.53.0 and langgraph 1.2.12 on tradefloor 0.8.6.

PydanticAI

TestModel answers every request with the arguments it is given. Pass your own model through the adapter's model= argument the same way.

import tradefloor as tf
from pydantic_ai import Agent
from pydantic_ai.models.test import TestModel
from tradefloor.integrations.pydantic_ai import PydanticAIAdapter

# a scripted model: no API key, the same answer at every decision
offline = TestModel(custom_output_args={
    "actions": [{"symbol": "AAA", "side": "BUY", "quantity": 200}],
    "rationale": "add to AAA"})

pm = Agent(instructions="You manage a small portfolio.")
adapter = PydanticAIAdapter(pm, model=offline)

market = tf.Universe.random(12, seed=4242)
card = tf.evaluate({"pm": adapter}, seed=4242, universe=market, days=5)["pm"]
print(card)
print(len(adapter.record), "decisions")
Scorecard('pm', pnl=2,055, return=+0.21%, trades=5, impact=+0.18bps, sharpe=n/a (short run), vol=0.7%, in_market=100%, exposure=0.03x)
5 decisions

For a live run, build the Agent with your provider's model, as PydanticAI's documentation describes, and leave out model=. Your agent's tools, deps and instructions are passed through unchanged, and the adapter binds the decision schema as the output type for each run.

OpenAI Agents SDK

ScriptedModel from agents.testing answers each call with a function of the input. payload_of reads back the observation the adapter sent.

import json

import tradefloor as tf
from agents import Agent
from agents.testing import ModelStep, ScriptedModel, assistant_message
from tradefloor.integrations.openai_agents import OpenAIAgentsAdapter, payload_of

def answer(call):
    # reads the same payload a real model is sent
    first = payload_of(call)["assets"][0]
    order = {"symbol": first["symbol"], "side": "BUY", "quantity": 200}
    return [assistant_message(json.dumps({"actions": [order],
                                          "rationale": "add"}))]

offline = ScriptedModel([ModelStep.respond(answer)] * 5)
pm = Agent(name="Portfolio Manager", instructions="You manage a small portfolio.")
adapter = OpenAIAgentsAdapter(pm, mode="live", model=offline)

market = tf.Universe.random(12, seed=4242)
card = tf.evaluate({"pm": adapter}, seed=4242, universe=market, days=5)["pm"]
print(card)
print(len(adapter.record), "decisions")
Scorecard('pm', pnl=2,055, return=+0.21%, trades=5, impact=+0.18bps, sharpe=n/a (short run), vol=0.7%, in_market=100%, exposure=0.03x)
5 decisions

OpenAIAgentsAdapter defaults to mode="replay", so a live run names mode="live", and openai_agent(agent) is the same adapter with live as its default. For a live run, give the Agent a model your account can use, set OPENAI_API_KEY and leave out model=. The adapter turns the SDK's tracing off for each run it starts unless you pass tracing=True, because tracing is on by default in the SDK and sends traces to OpenAI.

LangGraph

The adapter takes any compiled graph. By default it sends the payload under observation and as messages, and reads the decision from a decision key, an actions key or the last message.

from typing import Any, TypedDict

import tradefloor as tf
from langgraph.graph import END, START, StateGraph
from tradefloor.integrations.langgraph import LangGraphAdapter

class State(TypedDict, total=False):
    observation: dict[str, Any]
    decision: dict[str, Any]

def decide(state: State) -> State:
    # a rule in place of a model: buy 200 of the first name when flat
    first = state["observation"]["assets"][0]
    if first["position"] > 0:
        return {"decision": {"actions": [], "rationale": "holding"}}
    order = {"symbol": first["symbol"], "side": "BUY", "quantity": 200}
    return {"decision": {"actions": [order], "rationale": "open"}}

builder = StateGraph(State)
builder.add_node("decide", decide)
builder.add_edge(START, "decide")
builder.add_edge("decide", END)
adapter = LangGraphAdapter(builder.compile())

market = tf.Universe.random(12, seed=4242)
card = tf.evaluate({"graph": adapter}, seed=4242, universe=market, days=5)["graph"]
print(card)
print(len(adapter.record), "decisions")
Scorecard('graph', pnl=544, return=+0.05%, trades=1, impact=+0.04bps, sharpe=n/a (short run), vol=0.1%, in_market=100%, exposure=0.01x)
5 decisions

A node that calls a model makes the graph live, and nothing else changes. The graph decides five times and trades once, because after the first day it holds the name and answers with an empty action list, which is a decision to change nothing.

FinRobot

The FinRobot adapter predates the shared layer and is documented in the reference, FinRobot adapter. It needs the finrobot extra on Python 3.11.

Decisions and model calls

An adapter asks its framework once every every decision steps, and the default is 6. obs.step counts steps across the whole run, so with the default 6 steps a day that is one decision a simulated day, at each day's first step. With steps_per_day=3 the same default asks every second day. On the steps between decisions the adapter records the prices it saw and returns no orders.

One decision can cost several calls to the model provider, so a 20-day run is 20 decisions and can be many more provider calls:

  • A framework turn is one model call, and an agent that calls tools takes several turns per decision. The OpenAI Agents adapter caps a decision at max_turns=6, and the PydanticAI adapter at request_limit=8 requests.
  • PydanticAI retries a schema violation inside its own loop, and each retry is another request. The OpenAI Agents SDK 0.22 makes no client-side retry on a malformed answer.
  • A LangGraph graph makes whatever calls its nodes make.

The observation payload

The framework is shown the serialized payload and never the observation object. The payload is an allowlist written out field by field:

KeyWhat it holds
step, day, steps_per_dayThe clock. step counts the whole run.
macroThe published figures: the VIX, the policy rate, the corporate bond yield, inflation and the business-cycle phase as announced, late.
assetsPer name: the symbol, price, one-day and five-day returns and volatility from the prices the adapter has seen, the best bid and ask, average daily volume, max_order_shares (the participation cap), the position held, and any fundamentals you pass with fundamentals=.
portfolioCash, net worth, leverage, max_leverage, buying_power and the waiting limit orders.

It holds no fair value, no factor attribution and no macro path the run has not reached. That holds at every access level: the adapter's own act receives the read-only market view by default, or the live engine under trusted_agents=True, as What the agent sees sets out, and in both cases the framework receives the payload alone.

The decision is a list of actions, each with a symbol, a side (BUY, SELL, HOLD or CANCEL), a positive quantity in shares, and an optional limit_price that makes it a limit order. A bad action is refused on its own and listed in the scorecard's errors. An order over max_order_shares, 2% of the name's average daily volume by default, is cut to the cap, and an order below one share is dropped. Bringing an LLM agent has the full contract.

Reference

LLM adapters and MCP has every adapter's signature, the shared layer, error types, replay rules and the framework version each adapter was written against. When a run fails or makes no trades, Troubleshooting lists the messages and fixes.