A hard dollar limit per task run for AI tasks (using ctx.run.id) #4999
domondi1
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
AI tasks are where a single run can quietly get expensive: an agent loop that keeps calling tools, a fan-out over hundreds of items with
batchTriggerAndWait, or retries that re-run the whole thing.maxDurationbounds time, not spend. Here's one way to give each run its own dollar budget, keyed on the run id Trigger.dev already gives you.Run a small OpenAI-compatible gateway somewhere your tasks can reach (a VM or container next to your other services), with a shared token:
Then build the model inside the task, with the run's id and budget as headers:
Every call in the run reserves its worst-case cost before it's sent, so parallel calls inside one run can't all spend the same remaining money, and a call that doesn't fit is refused with a 402 before it reaches OpenAI (
APICallError,statusCode === 402). The run's budget is created the first time its id is seen and can't be raised afterwards, so a Trigger.dev retry of the same run shares the same budget instead of starting a fresh one. On the gateway host,inferrail work <run id>shows what that run cost.Caveats: Chat Completions only, set
maxOutputTokens(the reservation is based on it), the budget store is a SQLite file on the gateway host, so run one gateway rather than several replicas, and the model needs a price in Inferrail (gpt-4.1-mini,gpt-4.1,gpt-4o-miniare built in).Setup details: guide. I maintain Inferrail (open source, Apache-2.0), so take the suggestion with that in mind. Curious how others here bound the cost of a single AI run today.
All reactions