Local MCP server
tradefloor-mcp is an MCP server that runs on your machine and gives an MCP client, such as Claude Code or Claude Desktop, tradefloor's simulation and evaluation tools. A model uses it to study the simulator from outside: it builds universes and scenarios, scores strategies written as data, and explains why a price moved. No tool places an order, and no market lasts from one call to the next.
A model that trades runs through an LLM adapter on your machine, or on the hosted app, in beta.
Setup
Install the extra, which adds the mcp package, and register the server with your client. The server speaks MCP over stdio, so the client starts it as a subprocess.
pip install "tradefloor[mcp]" claude mcp add tradefloor -- tradefloor-mcp
The second line is for Claude Code. A client configured with a JSON file, as Claude Desktop is, takes the same command. Give the full path to tradefloor-mcp in the environment where you installed it, because the client does not start in your shell:
{
"mcpServers": {
"tradefloor": {
"command": "/path/to/venv/bin/tradefloor-mcp"
}
}
}A first tool call
In the client, ask the model to call describe_simulator. It returns what the simulator is, what it is measured to reproduce and what it cannot do, and the model should read it before anything else. A strategy is a JSON StrategySpec, never code, so the next step is to validate one and score it beside the baselines on one market.
The same calls from Python, through the MCP client library that the extra installs, show what a client receives. The block needs the extra, and was tested with mcp 2.3.0.
import asyncio
import json
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
SPEC = {"signal": {"kind": "momentum", "lookback_days": 1.0},
"portfolio": {"top_k": 5, "gross": 1.0}}
async def main():
server = StdioServerParameters(command="tradefloor-mcp")
async with stdio_client(server) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
tools = await session.list_tools()
print(len(tools.tools), "tools")
checked = await session.call_tool("validate_strategy", {"spec": SPEC})
print("valid:", json.loads(checked.content[0].text)["ok"])
scored = await session.call_tool(
"evaluate_strategies", {"strategies": {"mom": SPEC}, "days": 5})
result = json.loads(scored.content[0].text)
print("versus buy-and-hold:", result["versus_buy_and_hold"]["mom"])
asyncio.run(main())13 tools valid: True versus buy-and-hold: -18690.03
The strategy lost 18,690 against buy-and-hold on that one market, a 5-day run on seed 7. Every result carries a caveats list that says what it cannot show, such as a run too short for the model's validated scope, and a provenance block with the version, preset, seed and universe, so a model summarizing it sees the limits too. One seed is one sample: rank_strategies runs the same specs across seeds with a paired sign test.
The tools
| Tool | What it does |
|---|---|
describe_simulator | What the simulator is, what it is measured to reproduce and what it cannot do |
check_envelope | Whether a question is inside the validated scope, before anything runs |
validate_strategy | Checks a strategy spec and returns its fingerprint, without running it |
build_universe | Builds a universe: generated, concentrated in some sectors, or written by hand |
build_scenario | Composes a scenario and shows what it resolves to |
list_scenarios | The shipped scenarios, the constructors and the targets |
evaluate_strategies | Runs strategies and the five baselines on one market |
rank_strategies | Runs them across seeds, with a paired sign test between each pair |
run_stress_scenario | What a scenario does to each strategy, against the same market without it |
explain_price_move | A price move split into the eleven factors that sum to it |
explain | The random draws behind one day for one company |
start_job, check_job | Slow work in the background, up to 252 days |
A tool that runs a market returns the same result for the same arguments. A refused call comes back as a result with "ok": false and an error message, so the model can correct it. The scoring tools build a new market for each call and run the strategies in it under the read-only market view an ordinary agent gets, so a strategy cannot read fair value or fork the market. The oracle signal is the one exception, and its row says uses_hidden_state. No tool takes a preset or model coefficients. Every tool runs the default, pt-v20, and selecting another is a library call.
Limits
| Limit | Value |
|---|---|
| Days in a direct call | 60 |
| Days in a background job | 252 |
| Decision steps a day | 1 to 22, and days × steps at most 60 × 6 in a direct call |
| Names in a universe | 2 to 120 |
| Strategies in one call | 8 |
Seeds in rank_strategies | 2 to 12, six by default |
| Background jobs running at once | 2 |
A direct call over 60 days is refused, so pass the same arguments to start_job and poll check_job with the id it returns. Jobs live in the server process. A restart of the server, or of the client that started it, loses them and their results.
Reference
LLM adapters and MCP in the reference lists the tools and limits beside the LLM adapters, under the module tradefloor.mcp. Installation and extras covers install problems.