Putting a dollar cap on a single agent run #8919
domondi1
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
If you run Agent Framework agents for real users, one loop or a burst of tool calls can make a lot more model calls than you planned. As far as I can tell there's no per-run dollar limit built in yet, so here's one way to add it without changing the agent.
Point the Chat Completions client at a small local gateway and pass the run's id and budget as headers. We've been doing this with Inferrail:
Every call in that run reserves its estimated cost against the run's budget before it goes out, so parallel tool calls can't all spend the same remaining money. When a call doesn't fit, it fails with a 402 before it reaches the provider (you'll see a
ChatClientException), and you can end the run there.A few caveats: it's Chat Completions only for now, the reservation is an estimate so set
max_tokens, and only calls through the gateway count. Tested with agent-framework-core 1.19.0 and agent-framework-openai 1.14.4.Setup and the full snippet: recipe. I maintain Inferrail (open source, Apache-2.0).
All reactions