How to claim the monthly Claude Platform credit that now comes with Max and Team plans, what it pays for, what it never touches, and how to turn it into a live app before it expires
Since October 7, a Claude Max 20x subscription hands back its own price as API money: $200 a month.
That week Anthropic began rolling out a monthly Claude Platform credit to paying subscribers: $100 on Max 5x, $200 on Max 20x, and $20 or $100 per Team seat, pooled up to $500 a month - Claude Help Center. The money lands in a Claude Console organization that you link once, and it pays for the Claude API, the Message Batches API, the Console playground, Claude Managed Agents and the Claude Agent SDK. It arrived alongside Claude Haiku 5.5, a small model priced from $0.10 per million input tokens, and Anthropic halved Sonnet 5.5 cache reads to $0.10 the same day - Anthropic. Put the three together and a subscriber who has never written a line of API code now holds enough monthly budget to run the AI inside a small real product.
But most of this money will expire unused. The credit does not pay for interactive Claude Code, it does not raise your plan's usage limits, the organization you link cannot be changed without contacting support, and whatever you do not spend disappears at the end of each billing cycle. Many people pay for Max because of Claude Code, which is exactly the one product the credit excludes. To get anything from it you have to build something that calls the API, and you have to build it so that a leaked key or a runaway agent loop cannot spend through the credit into money you actually pay.
This guide covers how to claim the credit step by step, exactly what it covers (and the traps around Claude Code and API keys), what $100 buys on each current model, how to stretch it with caching, batches and the effort dial, a reference build that ships an app on it, the agent options (Agent SDK, Managed Agents and the new browser toolsets), what happens when it runs out, and how it compares with other routes from an idea to a live app. Every price and rule below was checked against Anthropic's own documentation and the other primary pages cited inline on October 10, 2026.
Contents
- What Anthropic Shipped on October 7
- Why a Subscription Now Comes With API Money
- How to Claim the Credit, Step by Step
- What the Credit Pays For, and What It Never Will
- What $100 Actually Buys, Model by Model
- Which Model to Spend It On
- Make the Credit Last: Caching, Batches and Effort
- Ship an App on the Credit: A Reference Build
- Agents on the Credit: Agent SDK, Managed Agents and Browser Use
- Routes From Idea to Live App, and Where the Credit Fits
- When the Credit Runs Out: Limits, Tiers and Failure Modes
- The Rest of the October Stack
- A Decision Framework
Where the Credit Can Go, Scored
Before the details, it helps to see every place a subscriber might try to spend the credit side by side, including the two places it cannot go. The table scores eight surfaces on the four things that decide whether the money turns into something useful: whether the credit pays for the surface at all, how much useful work each dollar buys, how quickly you get from nothing to a working feature, and how well the surface holds up once real users depend on it. Each cell carries the score and the fact behind it, and the sections that follow explain every row in depth.
The weighting reflects a simple priority. A surface the credit does not cover is worth little for this purpose no matter how good it is, so coverage carries 30%. Cost efficiency and production fit carry 25% each, because a credit that runs out in a week or a prototype that cannot serve users both waste the month. Time to first feature carries 20%, because the credit expires every cycle and a slow start burns part of it. Scores run from 0 (absent or unusable) to 10 (best available), and the final column is the weighted average.
| # | Surface | What It Does | Credit coverage (30%) | Cost per outcome (25%) | Time to first feature (20%) | Production fit (25%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | Messages API in your own app | Your server calls Claude per request | 10 - covered, first draw on the credit | 9 - Haiku 5.5 from $0.10/MTok, caching cuts repeats | 8 - one SDK call from a route handler | 9 - generally available, workspaces and spend limits | 9.1 |
| 2 | Message Batches API | Async jobs at half price | 10 - covered | 10 - 50% off input and output | 7 - results within 24 hours, most under 1 hour | 8 - 100,000 requests per batch, results kept 29 days | 8.9 |
| 3 | Claude Agent SDK (API key) | Claude Code's loop as a library | 10 - covered with a key from the linked org | 6 - agent loops re-read context every turn | 8 - a query() call in about 20 lines | 8 - you host it, Docker and cloud guides exist | 8.1 |
| 4 | Claude Managed Agents | Anthropic hosts the agent and sandbox | 10 - covered | 6 - tokens plus $0.08 per running session-hour | 8 - agent, environment, session, stream | 6 - beta, not eligible for ZDR or a HIPAA BAA | 7.6 |
| 5 | claude -p with a Console key | Headless Claude Code in scripts | 9 - covered when you run it with the org's key | 6 - same agent-loop token use | 9 - one command | 6 - good for cron and pipelines, not an app backend | 7.5 |
| 6 | Console playground | Test prompts in the browser | 10 - covered | 8 - standard rates, small volumes | 9 - no code, but nothing to deploy | 2 - no users, no endpoint | 7.3 |
| 7 | Interactive Claude Code | The terminal, IDE and desktop coder | 0 - not covered, uses plan limits | 8 - flat plan usage, no per-token bill | 9 - fastest way to write the app | 6 - builds production code, is not its runtime | 5.3 |
| 8 | Claude on Bedrock, Vertex AI or Foundry | Claude through a cloud provider | 0 - not covered | 7 - same list prices, regional endpoints +10% | 5 - cloud account and IAM setup first | 9 - enterprise cloud controls | 5.0 |
Read the table as a map rather than a verdict. The top two rows are where the credit does the most good for most subscribers: an app whose server calls the Messages API for interactive features, plus batch jobs for anything that can wait. The agent rows score lower on cost per outcome only because agents spend many tokens per task, not because they are a poor use of the money; Section 9 shows how to keep them inside the budget. The bottom two rows are there because they are the two mistakes people make most often in the first week: assuming the credit tops up Claude Code, and pointing a Bedrock or Vertex deployment at it.
Three caveats apply to every row. The scores assume you claim the credit into an organization you use for nothing else at first, so its balance is easy to watch. They assume current list prices, which moved again on October 7 and will move again. And they score the surface, not your project: a batch pipeline that nobody needs is still a waste of the month, however cheap each token is.
1. What Anthropic Shipped on October 7
October 7 looked like a model launch, and most coverage treated it as one. Anthropic released Claude Haiku 5.5, which it calls its cheapest, fastest and most capable small model, at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, rising to $0.50 and $2.50 above that threshold - Anthropic. Anthropic puts the average saving against Haiku 4.5 at about 75%, with requests under 100,000 tokens costing 90% less and longer ones 50% less. Haiku 5.5 is also the first Haiku-class model with the adjustable effort parameter, from low to max. We covered the model itself in our Haiku 5.5 launch analysis.
Two other changes shipped alongside it, and for anyone building an app they matter more than the headline model. First, Anthropic halved Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens, which it estimates cuts Sonnet 5.5 costs by about 20% on most agentic tasks. Second, it began rolling out the monthly API credit for Max and Team subscribers, worth $100 on Max 5x, $200 on Max 20x and up to $500 pooled on Team, usable on any Anthropic model through the Claude Platform. The same announcement added beta browser use and computer use support to the Python and TypeScript SDKs, so an app can drive a browser or a desktop through classes the SDK provides.
The credit details live in two places, and both are worth reading in full before you claim anything. The Help Center article sets out amounts, eligibility, the claim flow and the FAQ - Claude Help Center. The developer documentation adds the mechanics that matter once you build: how credits are ordered against purchased credits, how they interact with rate-limit tiers, and the exact error messages you will see when the balance hits zero - Claude Platform docs. Where the two overlap they agree, and this guide cites whichever states a rule more precisely.
The core terms fit in one table:
| Term | Max 5x | Max 20x | Team |
|---|---|---|---|
| Monthly credit | $100 | $200 | $20 per Standard seat, $100 per Premium seat |
| Cap | n/a | n/a | $500 pooled across the team |
| Refresh | Each billing cycle | Each billing cycle | Each billing month, monthly on annual plans |
| Rollover | None | None | None |
| Wait before claiming | 7 days on the plan | 7 days on the plan | 7 days on the plan |
| Who claims | The subscriber | The subscriber | Primary Owner or Owner |
The credit that came before it
The October credit is the second attempt at the same problem, and the first one explains why this one looks the way it does. On May 14, Anthropic announced that from June 15 the Agent SDK, claude -p and third-party tools built on them would stop drawing on subscription limits and would instead draw on a separate monthly credit: $20 on Pro, $100 on Max 5x, $200 on Max 20x - Zed. Tools that ran agents on a subscriber's plan, such as Zed's agent integration, would have moved onto metered billing. Coverage at the time described it as the end of flat-rate agent access on subscriptions - Tech Times. Anthropic paused the change on June 15, before it took effect.
That pause still stands. Anthropic's support page now says the Agent SDK, claude -p and third-party apps continue to count against subscription limits "for now", and that the previously announced Agent SDK credit is not available - Claude Help Center. The October credit is a different program: it sits in a Console organization instead of the subscription, it is billed at API rates against the Claude Platform, and Pro is no longer included. This matters because a lot of advice written between May and September describes the June design, not the one you can actually claim.
How to apply this section: treat October 7 as three separate things that happen to share a date. The credit gives you budget, Haiku 5.5 makes that budget stretch about ten times further than it would have on Haiku 4.5, and the cache cut lowers the price of the Sonnet 5.5 workloads that need more intelligence. The rest of this guide is about combining them, and the next section explains why Anthropic chose to hand subscribers API money at all.
2. Why a Subscription Now Comes With API Money
To use the credit well it helps to understand what problem it solves for Anthropic, because that tells you which uses it was designed for and which ones it will keep resisting. The structural fact is that a flat subscription and a metered API price the same underlying good, model inference, in two incompatible ways. A subscription works when usage is paced by a human: people type, read, think, and stop. The provider can predict the average, set limits on the peaks, and price the plan so that the average user is profitable. A metered API works for the opposite case, where a program decides how much to consume and the price has to follow consumption one token at a time.
Agents broke that division. Once Claude Code, the Agent SDK and headless claude -p could run for hours without a human in the loop, a subscription seat could be turned into an unattended workload, and the flat price stopped tracking the cost. Zed, which built one of the integrations affected, summarized a third-party estimate that subscriptions had been subsidizing heavy agent usage at roughly 15 to 30 times API pricing - Zed. Treat that range as an outside estimate rather than an Anthropic figure, but the direction is clear from Anthropic's own behavior: in May it tried to move automated usage onto metered billing, and in October it created a metered budget for it instead.
What the design tells you
Read as a design, the October credit is a bridge between the two pricing worlds. It keeps the subscription for what subscriptions are good at (interactive work in Claude, Claude Code and Cowork) and gives subscribers a metered allowance for what metered billing is good at (programs that call the model). The allowance is sized to the plan. On Max the credit equals the monthly price, $100 for Max 5x and $200 for Max 20x - Claude Help Center. On Team the per-seat credit equals the annual-billed seat price of $20 or $100 - Claude pricing. That symmetry looks deliberate: a subscriber gets the plan's price again as build money, without a single extra unit of interactive usage.
Three properties of the design are worth noticing because they shape strategy. The credit expires every cycle, so it rewards steady, recurring workloads over occasional big experiments. It is spent first, ahead of any credits you buy, so a project that outgrows it moves smoothly onto paid usage without a code change. And it does not count toward usage-tier advancement, so you cannot use free credit to climb into higher rate limits. Put plainly, it funds the period between "I wonder if this works" and "people are paying for this", which is exactly where a developer platform wants new builders to be.
The chart shows the symmetry and its one asymmetry. Max subscribers get the full monthly price back as credit. Team seats get the annual price back even when the team pays the higher monthly rate, and the pool stops at $500 however large the team grows. A five-person team with all Premium seats already reaches the cap, which means larger teams should think of the credit as a shared sandbox budget, not a per-person allowance.
Why this matters, and how to apply it
This matters because it tells you where the credit will stay generous and where it will not. Anthropic has a clear interest in subscribers building apps that call the API, since every app that outgrows $100 becomes a paying API customer, and a clear interest in keeping unattended agent workloads off flat subscriptions. Uses that fit the first interest (a product's runtime AI, background jobs, prototypes headed for real users) are what the credit is for. Uses that try to route around the second (piping the credit into an interactive coding tool, or reselling capacity) are outside its terms, and the Supplemental Credit Terms state that credits cannot be transferred or sold and are usable only by the holder of the account - Anthropic.
To apply this, plan the credit like a recurring budget rather than a windfall. Pick one workload that runs every week, size it to the credit with the math in Section 5, and let the first month tell you whether it is worth paying for once it outgrows $100. For a broader view of how inference costs should shape what you charge, our guide on pricing an AI product above its token costs works through the margin math.
3. How to Claim the Credit, Step by Step
Claiming takes about ten minutes, but one decision in the middle is close to permanent, so it is worth knowing the whole flow before you start. You link exactly one Claude Console organization to your plan. Each plan links to one organization, each organization can receive credits from only one plan, and you cannot change the link yourself; changing it means contacting support - Claude Platform docs. Whatever organization you pick is where your app's API keys must live for the credit to apply.
Before you begin, check eligibility and roles. Your Max or Team plan must be active, in good standing and at least seven days old. You can subscribe on the web, iOS or Android, but you claim on claude.ai in a web browser. On the plan side you need to be the Max subscriber, or a Primary Owner or Owner on Team; on the Console side you need the Owner, Admin or Billing role in the organization you link - Claude Help Center. Free, Pro and Enterprise plans are not eligible, while discounted Team plans for nonprofits and scientists are.
The flow itself has five steps:
- Open billing settings on claude.ai: Settings > Billing on Max, or Organization settings > Billing on Team.
- Find "API credits" and choose Link organization.
- Pick or create the Console organization that should receive the money.
- Accept the credit terms by confirming the link.
- Confirm the balance in the Console under Settings > Billing, in the Promotional credits section.
The fourth step is the irreversible one, so decide before you click. The safest default for most people is a new, dedicated organization used only for what you build on the credit. That keeps the balance easy to read, keeps any existing paid workloads in a different organization from surprise interactions, and lets you hand out keys without exposing a company account. If you already run a paid API organization for a product, linking that one instead means the credit offsets its bill automatically, which is the better choice when the product already has users. Anthropic also recommends signing in to the Console with the same email you use on claude.ai, and if linking fails it suggests trying again a few hours later, since the rollout runs over several days.
Create a workspace and a key
A Console organization starts with a Default Workspace that cannot be renamed, archived, deleted or given limits - Claude Platform docs. That last point matters: the only way to cap what one project can spend is to give it its own workspace and set that workspace's monthly spend limit. Every key and workspace in the organization draws from the same credit balance, so without a capped workspace, one leaked key can drain the month for every project at once.
The order that avoids trouble is to create a workspace named for the project (for example "Prod - support widget"), set its spend limit under Settings > Workspaces, then create the API key inside it. Spend limits can only be set lower than the organization's own limits, and rate limits per workspace are set on the Rate limits page. Every API response carries an anthropic-workspace-id header, which is the quickest way to confirm your key landed in the workspace you meant.
With the key created, a first request proves the whole chain works. This TypeScript snippet uses Haiku 5.5 with low effort and automatic prompt caching; the parameter names follow Anthropic's current docs for effort and caching - Claude Platform docs:
mkdir claude-credit-test && cd claude-credit-test
npm init -y && npm pkg set type=module
npm install @anthropic-ai/sdk
export ANTHROPIC_API_KEY="key-from-your-linked-org"
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const { data: message, response } = await client.messages
.create({
model: "claude-haiku-5-5",
max_tokens: 1024,
cache_control: { type: "ephemeral" },
output_config: { effort: "low" },
system: "You summarize customer emails in two sentences.",
messages: [{ role: "user", content: "Summarize: my order arrived late and the box was damaged." }]
})
.withResponse();
console.log(message.content);
console.log("workspace:", response.headers.get("anthropic-workspace-id"));
console.log("usage:", message.usage);
Run it with npx tsx and then open the Console's Cost page. The request should appear against your new workspace, and the Promotional credits balance should tick down by a fraction of a cent. If the request fails with a credit-balance error, either the credit has not arrived yet or the key belongs to a different organization from the one you linked; check the Promotional credits line first, then the key's organization.
The diagram is the mental model for everything that follows. Your plan feeds two pools that never mix. Interactive work in Claude and Claude Code draws on usage limits, exactly as it did before October 7. The credit lives only in the linked organization, reaches the model only through keys from that organization, and pays only for the API surfaces on the right. Why this matters: most of the questions in Anthropic's own FAQ for the credit come from mixing the two pools, either expecting the credit to top up Claude Code or expecting plan limits to cover an API call. How to apply it: before you spend a dollar, write down which pool each tool you use draws from.
4. What the Credit Pays For, and What It Never Will
Coverage is where the most expensive misunderstandings happen, because the excluded products are the ones subscribers use most. Anthropic's documentation lists the covered products explicitly, and anything outside the list should be assumed uncovered. The credit covers the Claude API (Messages and Message Batches), Claude Managed Agents, the Claude Agent SDK and the Console playground. It does not cover Claude Code, extra usage in the Claude apps, or Claude through cloud providers (Claude Platform on AWS, Amazon Bedrock, Google Cloud or Microsoft Foundry) - Claude Platform docs.
The reasoning behind the split follows from Section 2. Interactive Claude Code is a subscription product with its own usage limits, so paying for it from the credit would collapse the two pools back into one. Cloud-provider routes bill through the cloud provider, so Anthropic cannot apply a promotional balance held in its own Console. Everything that is covered is a surface where Anthropic meters tokens directly against your organization.
| Product | Covered | What actually pays |
|---|---|---|
| Claude API (Messages, Batches) | Yes | The credit, then purchased credits |
| Claude Managed Agents | Yes | The credit (tokens plus session runtime) |
| Claude Agent SDK with an org API key | Yes | The credit |
claude -p run by you with an org API key | Yes | The credit |
claude -p signed in with your plan | No | Plan usage limits |
| Interactive Claude Code (terminal, IDE, desktop, web) | No | Plan usage limits |
| Claude Code GitHub Action, IDE extension, desktop app runs | No | Counted as Claude Code usage |
| Bedrock, Vertex AI, Foundry, Claude Platform on AWS | No | The cloud provider's bill |
The Claude Code trap
The single most likely way to burn money by accident involves the API key itself. Claude Code checks for an ANTHROPIC_API_KEY environment variable, and if one is set it uses that key instead of your subscription, which produces API charges rather than drawing on your plan's included usage - Claude Help Center. Now imagine you followed Section 3, exported your new key in your shell to test your app, and then started Claude Code in the same terminal to keep building. Claude Code would pick up the key.
What happens next depends on the organization's balance. If it holds only the monthly credit, Claude Code shows "Credit balance too low", because the credit cannot pay for Claude Code sessions - Claude Platform docs. If you have also bought credits or turned on auto-reload, the session quietly bills those purchased credits at API rates. Either way you lose the flat-rate usage you are paying for. The fix is mechanical: keep the key in your app's environment file, never in your shell profile, and if you have exported it, unset ANTHROPIC_API_KEY before starting Claude Code. If Claude Code ever offers to continue on API credits when you hit a plan limit, decline unless you mean to pay API rates.
Headless runs and the GitHub Action
The rules for headless runs are subtle enough to state precisely. API credits cover claude -p and the Agent SDK when you run them yourself with an API key from your linked organization. If you run them signed in with your Claude plan, they use your plan's usage limits instead. And runs started by the Claude Code GitHub Action, an IDE extension or the Claude desktop app count as Claude Code usage, so the credit does not cover them even when they use -p - Claude Help Center.
In practice this splits automation into two camps. A cron job or a pipeline script that you own, calling claude -p --output-format json with the organization's key in its environment, is a covered workload (Anthropic documents the headless mode for exactly this kind of scripting - Claude Code docs), and a reasonable way to spend part of the credit on recurring chores such as nightly dependency reviews. A pull-request bot running through the official GitHub Action is not covered, however it authenticates. If you want a covered code-review bot, build it on the Agent SDK in your own runner rather than on the Action. For running Claude Code unattended on your plan instead, see our guide to Claude Code's auto mode.
What you may not build with it
One more rule shapes product design rather than billing. Anthropic does not allow third-party developers to offer claude.ai login or Claude plan rate limits inside their own products unless Anthropic has approved it, and this explicitly includes agents built on the Agent SDK; products must authenticate with API keys - Claude Code docs. So you cannot build an app that lets your users sign in with their Claude subscription and spend their plan, the way OpenAI now allows with its own login (our Sign in with ChatGPT guide covers that model). On Claude, the developer pays for the app's inference, and the October credit is Anthropic's way of covering the first $100 or $200 of it.
Use of the Agent SDK in products for your own customers falls under Anthropic's Commercial Terms, and the SDK's branding rules let you say "Powered by Claude" but not present your product as Claude Code. Why this matters: a design that depends on users' Claude plans is not allowed without Anthropic's prior approval, so plan your pricing around paying for inference yourself. How to apply it: budget your app's AI as your cost from day one, use the credit to absorb it while you find out whether users want the product, and price above it before the credit runs out.
5. What $100 Actually Buys, Model by Model
"A hundred dollars of API credit" means nothing until you convert it into the units your app consumes. Models bill per million tokens (MTok), separately for input and output, with discounts for cached input and batched requests. A token is a fragment of text; on Claude models from the 4.7 generation onward, Anthropic's newer tokenizer produces roughly 30% more tokens for the same text than older models did - Claude Platform docs. Thinking tokens, which current models use by default, bill as output. So the honest unit is not "messages" but "tokens per request", measured on your own prompts.
Here are the current list prices for the four models a builder is likely to use, verified on Anthropic's pricing page on October 10:
| Model | Input /MTok | Output /MTok | Cache read /MTok | Batch input / output |
|---|---|---|---|---|
| Claude Haiku 5.5 (prompts up to 100K) | $0.10 | $0.50 | $0.01 | $0.05 / $0.25 |
| Claude Haiku 5.5 (prompts over 100K) | $0.50 | $2.50 | $0.05 | $0.25 / $1.25 |
| Claude Sonnet 5.5 | $2 | $10 | $0.10 | $1 / $5 |
| Claude Opus 5.5 | $4 | $20 | $0.20 | $2 / $10 |
| Claude Fable 5.1 | $10 | $50 | $0.25 | $5 / $25 |
Raw, the $100 Max 5x credit buys about 1 billion Haiku 5.5 input tokens or 200 million output tokens; 50 million Sonnet 5.5 input tokens or 10 million output tokens; half that on Opus 5.5; and 10 million input or 2 million output tokens on Fable 5.1. Those numbers are too abstract to plan with, so convert them into requests. Take a typical app request of 2,000 input tokens and 500 output tokens: a system prompt, some user context, and a short answer. The chart below prices 1,000 of those requests on each model.
At those rates the $100 credit covers roughly 222,000 Haiku 5.5 requests, 11,100 on Sonnet 5.5, 5,550 on Opus 5.5 and 2,200 on Fable 5.1. The hundredfold gap between the cheapest and the most capable model is the most important number in this guide, because it means model choice moves your budget far more than any prompt tweak. It also explains why Haiku 5.5 changes what the credit can do: on Haiku 4.5 the same $100 would have covered about 22,000 of these requests.
Output size dominates on long generations. Anthropic's own Sonnet 5.5 launch page shows models writing complete interactive programs in a single HTML file, and the frame below, from that page's side-by-side demos, shows one run that produced 4,158 output tokens. At Sonnet 5.5's $10 per million output tokens, that whole program would cost about four cents to generate, which is the right intuition for code generation: individual artifacts are cheap, and volume is what spends the budget.
Workloads the credit can carry
Request counts become useful when you map them onto workloads you might actually run. The estimates below use list prices and stated assumptions; treat them as starting points to replace with your own measured token counts after a week of traffic. Each one is sized against the $100 Max 5x credit, so double them for Max 20x.
| Workload | Model and setup | Assumptions | Monthly cost | Fits in $100? |
|---|---|---|---|---|
| Support chat on your site | Haiku 5.5, cached system prompt | 30 messages per user, 3,000 in / 400 out each, 2,500 cached | about $0.008 per user | Up to about 12,000 users |
| Nightly document digest | Haiku 5.5, Batches API | 10,000 docs of 5,000 tokens, 300-token summaries, 30 nights | about $97.50 | Just |
| Coding agent sessions | Sonnet 5.5, Managed Agents | 1 hour, 50,000 in / 15,000 out per session | about $0.33 per session | About 300 sessions |
| In-app writing assistant | Sonnet 5.5, cached prompt | 30 uses per user, same token profile as support chat | about $0.16 per user | Up to about 635 users |
| Research reports | Opus 5.5 at medium effort | 20,000 in / 4,000 out per report | about $0.16 per report | About 625 reports |
The pattern in the table is the practical lesson: high-volume, short-answer features belong on Haiku 5.5, where even thousands of users fit inside the credit, while Sonnet and Opus fit when each request is valuable enough that hundreds per month is the right volume. The digest row also shows how batching changes the math, since the same job run synchronously would cost twice as much and blow through the month in about two weeks. For a deeper look at matching each request to the cheapest model that can handle it, see our guide to cutting agent costs with model routing.
How to apply this section: before building anything, write your feature's expected token profile (input, output, how much of the input repeats) and multiply it out for your expected users. If the result is under half the credit, you have room to experiment. If it is above the credit, decide now whether caching, batching, a cheaper model or a usage cap brings it under, because those choices are far cheaper to make before launch than after.
6. Which Model to Spend It On
Model choice is the biggest lever on the credit, and Anthropic's own guidance is more specific than most commentary. Its documentation suggests starting with Claude Opus 5.5 for most workloads, using Claude Fable 5.1 for demanding reasoning and long-horizon agent work or when Opus 5.5 at higher effort still falls short, using Sonnet 5.5 where you want the best combination of speed and intelligence, and using Haiku 5.5 for high-volume, latency-sensitive tasks such as classification, extraction and routing - Claude Platform docs. All four have a 1 million token context window and 128,000 tokens of maximum output, and all four use adaptive thinking steered by effort.
Those recommendations are written for quality, not for a fixed monthly budget, so a credit holder should read them slightly differently. When the budget is $100, start each feature on the cheapest model that might work, measure, and move up only where quality fails. That usually means Haiku 5.5 for anything a user triggers frequently, Sonnet 5.5 for features where the answer's quality is the product, and Opus 5.5 or Fable 5.1 reserved for low-volume, high-value work such as planning an agent's approach or reviewing its output.
How good is the cheap model?
The case for Haiku 5.5 rests on Anthropic's benchmark table, which compares it with Haiku 4.5 and OpenAI's small GPT-6 Luna. Vendor benchmarks are run by the vendor and should be read as directional, but the size of the gaps is large enough to matter. On computer-use tasks, terminal-based coding and chart understanding, Haiku 5.5 scores several times higher than the model it replaces and well above GPT-6 Luna.
Anthropic is explicit about the limits, too. It says Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding, and positions Haiku 5.5 for narrowly scoped jobs such as summarization, compaction and subagent work. Its own Terminal-Bench 4.0 number makes the point: Haiku 5.5 scores 39.2% where Sonnet 5.5 scores 70.6%. So the right reading is not "use Haiku for everything" but "use Haiku for every request that does not need Sonnet", which in most apps is the majority of requests.
Sonnet 5.5 is the model most credit holders will reach for when quality matters, and Anthropic's launch film is a short, useful look at what it is designed for. In it, Anthropic describes Sonnet 5.5 as more than 30% faster than Sonnet 5, with clearer writing, and strongest at well-scoped everyday tasks, bug fixes and polished documents, while Opus 5.5 handles work that needs more careful judgment. The launch page states the same speed claim and lists the pricing - Anthropic.
The video's framing maps neatly onto credit strategy. If a feature is "well-scoped and everyday" (rewrite this, extract that, answer from these documents), Sonnet 5.5 or Haiku 5.5 is the right tier and the credit goes far. If a feature needs "careful judgment" across many steps, Opus 5.5 earns its higher price per request, but you should expect hundreds rather than thousands of runs per month inside the credit. For a direct comparison of the previous generation's two middle tiers, our Opus 5 vs Sonnet 5 guide explains the trade-off in more depth, and our Claude 5.5 migration guide covers the breaking changes if you are moving older code onto these models.
The effort dial is a second price list
Every current model accepts an effort setting (low, medium, high, xhigh, max) that changes how many tokens it spends thinking, calling tools and explaining. Defaults differ: Opus 5.5 and Haiku 5.5 default to medium, Sonnet 5.5 and Fable 5.1 to high - Claude Platform docs. Because thinking bills as output, effort can change a request's cost severalfold without touching the model or the prompt.
Anthropic's per-model advice is concrete. For Haiku 5.5 it recommends starting at medium, dropping to low for chat, short tool tasks and simple high-volume requests, and noting that at low effort the model is more likely to skip a search or stop early in long agent prompts. For Sonnet 5.5 it recommends medium for well-specified agentic coding and medium or low for chat. Why this matters for the credit: an app left on Sonnet 5.5's default high effort for simple chat can spend far more than the same app at medium with no visible quality loss. How to apply it: set effort explicitly on every request, run your own small evaluation at two levels, and keep the cheaper one where results hold. Our guide to setting the effort dial walks through how to run that sweep.
7. Make the Credit Last: Caching, Batches and Effort
Once a model is chosen, three platform features decide whether the credit lasts the month: prompt caching, the Batches API and the effort setting from the previous section. They stack, and together they can cut the cost of a typical app's AI by more than half without changing what users see. The reason they work is structural. Most app requests repeat a large fixed prefix (instructions, policies, tool definitions, reference documents) and vary only in a small user-specific tail, and most background work does not need an answer within seconds. Caching stops you paying full price for the repeated prefix, and batching stops you paying full price for urgency you do not need.
Caching is the bigger lever for interactive apps. A cache read costs a small fraction of normal input: on Sonnet 5.5 and Opus 5.5 it is 5% of the base input price, on Fable 5.1 2.5%, and on Haiku 5.5 and most other models 10% - Claude Platform docs. Writing to the 5-minute cache costs 1.25 times the input price and to the 1-hour cache twice the input price, so caching pays for itself after one read in the short window or two in the long one. Cached tokens also do not count toward your input-tokens-per-minute rate limit on current models, so caching raises your effective throughput as well as lowering cost.
Turn caching on
The simplest form is automatic caching: add one cache_control field at the top level of the request and the API manages cache breakpoints as the conversation grows. The minimum cacheable prompt on Haiku 5.5, Sonnet 5.5 and Opus 5.5 is 512 tokens; anything shorter is processed without caching and without an error, so check cache_read_input_tokens in the response to confirm hits - Claude Platform docs. Here is the pattern in Python with a long, stable system prompt:
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the linked org
POLICY = open("support_policy.md").read() # a few thousand stable tokens
def answer(question: str) -> str:
response = client.messages.create(
model="claude-haiku-5-5",
max_tokens=800,
cache_control={"type": "ephemeral"}, # 5-minute cache; use "ttl": "1h" for slow traffic
output_config={"effort": "low"},
system=POLICY,
messages= [{"role": "user", "content": question}],
)
print(response.usage) # watch cache_creation_input_tokens vs cache_read_input_tokens
return "".join(b.text for b in response.content if b.type == "text")
The most common way to lose the cache is changing something early in the prompt on every request, such as a timestamp or the user's name inside the system prompt. Keep everything that varies at the end, after the stable prefix. Changing the top-level effort between requests also restarts the cache, so pick one level per conversation. Prompt caches are isolated per workspace on the Claude API, which means a separate workspace per project (good for spend limits) also means separate caches. Our full guide to prompt caching covers the less obvious ways teams lose cache hits.
Batch everything that can wait
The Message Batches API charges 50% of standard prices for input and output, accepts up to 100,000 requests or 256 MB per batch, finishes most batches within an hour, and keeps results for 29 days - Claude Platform docs. Batches can include tool use, web search and most beta features, and their discounts stack with caching, though cache hits inside a batch are best-effort. Anything your app does on a schedule (digests, classification of yesterday's records, enrichment, moderation of new content, regenerating help articles) belongs here:
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request
batch = client.messages.batches.create(
requests= [
Request(
custom_id=f"doc-{doc_id}",
params=MessageCreateParamsNonStreaming(
model="claude-haiku-5-5",
max_tokens=400,
messages= [{"role": "user", "content": f"Summarize in 3 bullets:\n\n{text}"}],
),
)
for doc_id, text in documents_to_summarize()
]
)
print(batch.id, batch.processing_status) # poll until "ended", then stream results
Results come back in any order, so match them by custom_id, and validate one request shape against the normal Messages API first, since batch validation errors only appear once the whole batch has finished. One caveat matters for budget control: because batches process concurrently, Anthropic notes they can slightly overshoot a workspace's spend limit.
What caching does to your monthly bill
The chart below makes the combined effect concrete for an app with a chat-style feature. It assumes each active user sends 30 requests a month, each with 3,000 input tokens (2,500 of them a stable, cached system prompt) and 400 output tokens including thinking, at list prices. Cache-write costs are left out because with steady traffic they are small next to the reads.
Read the crossover points against your credit. Uncached Sonnet 5.5 exhausts $100 at about 333 active users and $200 at about 667. Caching the prompt nearly doubles that, to about 635 and 1,270. Cached Haiku 5.5 carries about 12,100 users on $100 and more than 24,000 on $200. Why this matters: the difference between "the credit covers my beta" and "the credit covers my first ten thousand users" is two decisions you can make in an afternoon. How to apply it: put the stable prefix first and cache it, send every non-urgent job through batches, set effort explicitly, and route each request to the cheapest model that passes your quality check.
8. Ship an App on the Credit: A Reference Build
Everything so far is preparation; the credit only pays off when something real calls it. This section walks through a reference build that a single person can ship in a weekend: a web app whose server calls Claude for one core feature, stores its data in a managed database, and deploys to a host with a free tier. The specific stack is an example, not a requirement. What matters is the shape, which keeps the API key on the server, isolates the project's spend, and makes the expensive decisions (model, effort, caching) in one place.
The example below uses Next.js 16.4, released October 6, which now recommends its Cache Components model for every app and enables it by default in new projects created with create-next-app - Next.js. Any framework with server-side routes works the same way. For the database and hosting choices, our rankings of databases for your product and where to deploy your app compare the options in depth.
The build, in order
The order of work matters more than the tools, because the two failure modes for a credit-funded app are a leaked key and an unbounded loop. Both are prevented by decisions you make before the first deploy. Start from the workspace and limits you created in Section 3, then build outward from the server.
The guiding rule is that the browser never talks to Claude directly. Every call goes through a server route you control, which is the only place the API key exists and the only place you can enforce who may call the model and how often. That one constraint turns the credit from a shared pot anyone with your page's source could drain into a budget that only your own code can spend. It also gives you a single spot to change models, add caching or log usage later, without touching the front end. With that in mind, the build runs in five steps.
- Scaffold the app with
npx create-next-app@latestand add the SDK withnpm install @anthropic-ai/sdk. - Put the key in server env, in the host's environment settings and a local
.env.localthat is never committed. - Write one route handler that validates input, calls Claude, and returns the answer.
- Add per-user limits in your own code, such as a daily request count stored in the database.
- Deploy, then watch the Console's Cost page for a day before telling anyone.
The fourth step is the one most tutorials skip and the one that protects the credit. A workspace spend limit stops the bleeding when something goes wrong, but it stops it for every user at once. A per-user limit in your own code stops one abusive or buggy client without taking the feature down for everyone. Together they give you two lines of defense: one inside your product, one at Anthropic's edge. Our pre-launch security checklist covers the rest of what an AI-built app needs before real users touch it.
Here is a complete route handler for a "summarize this text" feature. It keeps the key on the server, caps input size, caches the stable instructions and sets effort explicitly:
// app/api/summarize/route.ts
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic(); // ANTHROPIC_API_KEY from the server environment
const INSTRUCTIONS = `You summarize text for busy readers.
Return three short bullet points and nothing else.`;
export async function POST(req: Request) {
const { text } = await req.json();
if (typeof text !== "string" || text.length === 0 || text.length > 20000) {
return Response.json({ error: "Send between 1 and 20,000 characters." }, { status: 400 });
}
// checkAndCountDailyLimit is your own function backed by your database
// if (!(await checkAndCountDailyLimit(req))) return Response.json({ error: "Daily limit reached." }, { status: 429 });
const message = await client.messages.create({
model: "claude-haiku-5-5",
max_tokens: 600,
cache_control: { type: "ephemeral" },
output_config: { effort: "low" },
system: INSTRUCTIONS,
messages: [{ role: "user", content: text }]
});
const summary = message.content
.filter((block) => block.type === "text")
.map((block) => block.text)
.join("");
return Response.json({ summary, usage: message.usage });
}
Returning usage during development is deliberate: it lets you see real token counts per request in the browser's network panel, which is the data you need to redo Section 5's math with your own numbers instead of assumptions. Remove it, or log it server-side, before launch. The cached instructions here are short, below the 512-token caching minimum, so in a real app you would move longer stable material (style rules, examples, product facts) into the system prompt where caching starts to pay.
Why this shape, and how to grow it
The shape matters because it decides what happens at month two. An app built this way can switch models with one string, move background work to batches without touching the interactive path, and move from credit to purchased credits with no code change at all, since included credits are spent before purchased ones. It also keeps your options open on pricing: once you know your cost per active user, you can set a price that covers it, which our guide to metered billing for AI products explains how to implement.
How to apply it: ship the smallest version with one Claude-powered feature, a per-user limit and a capped workspace, then spend the rest of the month learning from real usage. If you want a fuller walkthrough of building and deploying with Claude Code doing most of the typing, our guide to building a live app with Claude Code covers the workflow end to end, and Claude Code itself runs on your plan's usage limits, not on the credit. A common first feature for a credit-funded app is an assistant that answers questions about your own product, which our guide to building a support agent for your site covers in detail.
9. Agents on the Credit: Agent SDK, Managed Agents and Browser Use
Agents are the most interesting use of the credit and the fastest way to spend it. An agent is a program that plans its own steps and calls tools in a loop until a task is done, and every turn of that loop re-sends context to the model. That is why the scoring table rated agent surfaces lower on cost per outcome: a task that a single request could answer in 3,000 tokens might take an agent 100,000 tokens across twenty turns. In exchange, agents can do work no single request can, such as fixing code across a repository, researching across many pages, or operating software on a user's behalf.
The credit covers three ways to run agents, and they differ mainly in who operates the loop and the sandbox. The Claude Agent SDK gives you Claude Code's agent loop, tools and context management as a Python or TypeScript library that runs in a process you operate. Claude Managed Agents lets Anthropic host the loop and a sandbox for you, configured through the API. And the new browser and computer toolsets in the client SDKs let your own loop drive a browser or a desktop that you supply - Claude Code docs.
The Agent SDK
The Agent SDK is the closest thing to "Claude Code inside your product". Installation is one package, @anthropic-ai/claude-agent-sdk for TypeScript or claude-agent-sdk for Python, and both bundle the Claude Code binary. It reads ANTHROPIC_API_KEY from the environment of the process that runs it and does not load .env files automatically - Claude Code docs. With a key from your linked organization, its usage draws on the credit. A minimal agent that reviews and fixes a file looks like this:
import { query } from "@anthropic-ai/claude-agent-sdk";
for await (const message of query({
prompt: "Review utils.py for bugs that would cause crashes. Fix any issues you find.",
options: {
allowedTools: ["Read", "Edit", "Glob"],
permissionMode: "acceptEdits"
}
})) {
if (message.type === "result") console.log(`Done: ${message.subtype}`);
}
Two settings keep an SDK agent inside the budget. Restrict allowedTools to what the task needs, since every tool definition adds input tokens to every turn, and give the agent a clear stopping condition in the prompt. Run it in its own workspace with a spend limit, because an agent stuck in a retry loop is one of the fastest ways to empty a month's credit in an afternoon. For coordinating several agents at once, our parallel agents playbook covers the patterns that keep costs predictable.
Managed Agents
Managed Agents moves the loop and the sandbox to Anthropic. You define an agent (model, system prompt, tools, MCP servers, skills), an environment (an Anthropic-managed cloud sandbox or a self-hosted one), and start sessions that stream events back. It is in beta behind the managed-agents-2026-04-01 header, which the SDKs set for you, and it is enabled by default for API accounts - Claude Platform docs. Because sessions are stateful and stored server-side, Managed Agents is not currently eligible for Zero Data Retention or a HIPAA BAA, which rules it out for some regulated data.
Pricing has two parts: tokens at normal model rates, plus $0.08 per session-hour of runtime, metered only while a session is running, not while it waits for your next message - Claude Platform docs. The Batch discount does not apply. Anthropic's quickstart creates an agent, an environment and a session in a few calls:
from anthropic import Anthropic
client = Anthropic()
agent = client.beta.agents.create(
name="Research Assistant",
model="claude-sonnet-5-5",
system="You research questions and write short, sourced answers.",
tools= [{"type": "agent_toolset_20260401"}],
)
environment = client.beta.environments.create(
name="research-env",
config={"type": "cloud", "networking": {"type": "limited", "allow_package_managers": True}},
)
session = client.beta.sessions.create(agent=agent.id, environment_id=environment.id, title="First run")
print(session.id) # then open an event stream and send a user.message event
Managed Agents fits best when you want agent behavior without operating containers: long-running research, document processing, or a coding assistant your users can trigger from a web page. Anthropic's quickstarts include complete chat apps built on Vercel's Chat SDK, assistant-ui and CopilotKit - Claude Platform docs. The source for those apps sits in Anthropic's public quickstarts repository - GitHub. Note that limited networking with an empty allowed_hosts list also blocks web search and fetch results, so list the hosts your agent needs.
Browser and computer use
The October SDK release added beta classes for the browser use and computer use tools in Python and TypeScript. You subclass BetaAbstractBrowserToolset20260801 (or the computer equivalent) and implement one method per action, such as navigate, screenshot or left_click, against your own browser automation; the SDK runs the loop, applies the policies you pass and builds each tool result - Claude Platform docs. The SDK does not include a browser, a desktop or a URL policy. Browser Use, Browserbase, Daytona and E2B publish their own integrations if you would rather not run the browser yourself.
Budget for the overhead. Declaring the browser toolset with its default members adds about 6,600 input tokens to every request, and the computer toolset about 4,500, before a single screenshot; screenshots then bill as image input on every turn that includes them - Claude Platform docs. Caching the toolset definition helps, but as a rough estimate a 20-step browser task on Sonnet 5.5 can still cost anywhere from tens of cents to about a dollar depending on screenshots and cache hits, so measure before you promise users unlimited runs. The safety guidance is unusually blunt: write a URL policy, put egress rules on the container, block the cloud metadata address, and never give the agent a browser profile signed in to anything you would not hand to Claude.
Anthropic's own short film on working with its newest models makes a point that applies directly to agents on a budget: people tend to give capable models tasks that are too small and too tightly scripted, which wastes turns. The video walks through three habits that hold Opus 5.5 and Sonnet 5.5 back, and what to do instead.
The budget lesson from the video is counterintuitive: a well-specified, larger task often costs fewer tokens than the same work split into many small, micromanaged requests, because each small request repeats context and instructions. Why this matters: agent spending is dominated by the number of turns, not the price per token. How to apply it: give an agent one complete goal with clear acceptance criteria, cap its tools and turns, run it in a capped workspace, and compare cost per finished task rather than cost per request. If your agent needs to act in the world with real money or accounts, our guide to agent identities instead of shared API keys covers how to scope what it can do.
10. Routes From Idea to Live App, and Where the Credit Fits
The credit pays for a product's AI, but someone still has to build the product, and the route you choose decides which parts of the bill the credit touches. This is the point most first-week advice misses. Building and running are two different costs. Building is the work of writing and changing code, which for most Max subscribers happens in Claude Code on their plan's usage limits. Running is the AI your finished product calls when users touch it, which is what the credit pays for. A good plan uses the plan for building and the credit for running, and does not confuse the two.
Different routes split those costs differently. Some keep everything under your control and on your accounts, at the price of assembling the stack yourself. Some hand the building to a hosted tool that bills its own credits. Some sit in between. None is universally right; the choice depends on whether you want to own a codebase, how fast you need something live, and how much of the product's ongoing development you want to do by hand. The comparison below describes what each route actually charges for, so you can see where the credit lands.
| Route | Who builds | What the API credit pays | What you pay besides | Fits best when |
|---|---|---|---|---|
| Claude Code + your own repo and host | You, with Claude Code on your plan | The app's runtime Claude calls | Your plan, hosting, database | You want to own and understand every file |
| Agent SDK or Managed Agents app | You write the harness | The agent's tokens (and session-hours) | Hosting, or nothing extra on Managed Agents | The product is the agent itself |
| Hosted AI app builders | The builder, on its own models | Nothing for building; runtime only if the app uses your own Anthropic key | The builder's plan and credits | You want a prototype in an hour |
| Founden desktop app | An AI builder running Claude Code on your Mac, on your plan | The product's AI features, when your own Anthropic key is in its workspace | Founden credits; builds run through Founden's backend | You want the product, site and back office built as one company |
The first row is the default for developers and the route Sections 3 and 8 describe. Its strength is ownership: the code, the accounts and the bill are all yours, and the credit applies to the app's API calls with no intermediary. Its cost is assembly time, since you choose and wire the framework, database, auth, payments and hosting yourself. The second row is the right choice when the agent is the product, for example a research assistant or a coding helper your users trigger.
Hosted app builders sit at the other end. They make a working prototype fast, but their own AI usage is billed in their own credits, so your Claude credit cannot pay for the building, and it only pays for the finished app's AI if the tool lets you supply your own Anthropic key for runtime calls. Our comparison of Lovable, v0 and Bolt by cost per app covers what those credits cost and when to graduate from a vibe-coding tool to your own codebase.
Founden is one of the routes in between. Its desktop app runs the Claude Code CLI (or OpenAI's Codex) on the founder's own Mac, billed to the founder's own Claude or ChatGPT plan, and keeps the company's source as a git repository on that Mac; a founder's own AI keys in the company's workspace are the ones its product uses in production. In credit terms, the building draws on your Max plan's usage limits like any other Claude Code session, and the credit can pay for the finished product's Claude calls through your own key. The trade-offs are that builds run through Founden's backend and the desktop app is macOS only.
Why this matters: the credit is most valuable on routes where your own Anthropic key sits in the runtime path, because only then does every user request draw on money you already received. How to apply it: whichever route you pick, check where the product's AI calls are billed. If the answer is "my linked organization", the credit is working for you; if it is "the tool's account", you are paying twice, once in your plan and once in the tool's credits.
11. When the Credit Runs Out: Limits, Tiers and Failure Modes
Every credit-funded project eventually meets the edge of the balance, and the behavior at that edge is well defined. Included credits are spent first, then any credits you have purchased. If the organization has no purchased credits or auto-reload, API requests simply stop until the next monthly grant, and usage is never charged to your Claude plan. Organizations invoiced through Anthropic's sales team are billed normally beyond the credit - Claude Platform docs. The error your app receives reads: "Your credit balance is too low to access the Anthropic API. Please go to Plans & Billing to upgrade or purchase credits."
That stop is a safety feature as much as a limit. A credit-only organization cannot surprise you with a bill, which makes it a good home for experiments and agents you do not fully trust yet. The flip side is that a live product on a credit-only organization goes dark mid-month when the balance runs out, so any app with real users needs either purchased credits behind the credit or a graceful degradation path, such as falling back to a cheaper model or showing a "busy" message, when the API returns a balance error.
Tiers and spend caps
The credit also interacts with rate-limit tiers, which matters once an app has traffic. Organizations linked to an eligible plan move up to at least the Start tier, and organizations whose monthly credit is over $200 move up to at least the Build tier; otherwise claiming does not change your tier, and the credits do not count toward moving up - Claude Platform docs. Because the rule says "over $200", a Max 20x subscriber's exactly-$200 credit leaves the organization at Start, while a Team pool of $260 or more reaches Build, a quirk that MIXED's coverage also flagged - MIXED.
The tiers carry monthly spend caps of $500 (Start), $1,000 (Build) and $200,000 (Scale), and usage paid by the credit counts toward the cap - Claude Platform docs. At Start, Sonnet 5.5, Opus 5.5 and Haiku 5.5 each allow 1,000 requests per minute, 2 million input tokens per minute and 400,000 output tokens per minute, which is ample for an early app, especially since cached input does not count toward the input limit. Two timing details trip people up: spend caps reset at 00:00 UTC on the first of each calendar month, while the credit refreshes on your billing cycle, so the two windows rarely line up; and hitting the tier cap returns HTTP 429 with the code enforced_spend_limit_reached and no retry-after header, so automatic retries will keep failing until the month turns.
Failure modes worth preventing
Most credit losses in the first month come from a short list of avoidable mistakes, and each one has a cheap, specific prevention. The first two happen before you write any product code. Linking the wrong organization cannot be undone without support, so decide which organization gets the credit before you confirm the link, not after. Exporting the key in your shell lets Claude Code pick it up and bill it instead of your plan, so keep keys in app environment files and run unset ANTHROPIC_API_KEY before starting Claude Code in any terminal where you tested your app.
The next two happen once code exists. Shipping the key to the browser, for example by calling Claude from client-side JavaScript, lets anyone who opens your page copy it and spend your balance, so call Claude only from server code. Unbounded agent loops, where an agent retries a failing step until the balance is gone, are stopped by a workspace spend limit at Anthropic's edge plus a turn cap in your own code, since the spend limit alone will also stop every other feature in that workspace.
The fifth mistake is about billing settings. It is tempting to assume that auto-reload waits until the monthly credit is used up, but Anthropic's docs say it watches only your purchased balance and reloads when that balance reaches its threshold, even if included credits remain. Turning it on with a high threshold can therefore buy credits you did not need yet; a low threshold, or leaving auto-reload off until the product has paying users, avoids that. A sixth risk is organizational rather than technical: on Team plans, seats added count from the next billing month's credit, and removed seats lower the pool after renewal, so a team that shrinks may find its pool smaller than it budgeted.
Plan changes follow clear rules. Cancelling or downgrading stops new credits, but credits already granted stay usable until they expire. Upgrading from Max 5x to Max 20x grants prorated credits immediately and then $200 per cycle. Moving from Max to Team ends the Max link, and a Team owner can claim the team's credits after seven days on the plan - Claude Help Center. Why this matters: a credit that fails silently mid-month is worse than no credit, because users find out before you do. How to apply it: alert on balance, handle the balance error in your app, and keep a small purchased balance behind any product with real users.
12. The Rest of the October Stack
The credit landed in a week when much of the surrounding toolchain moved too, and several of those changes make a credit-funded app easier to build or cheaper to run. None of them is required, but each removes a piece of work a solo builder would otherwise do by hand. The common thread is that infrastructure vendors are now designing for agents as first-class users: agents that upgrade frameworks, buy domains, create databases on demand and pay for resources per request.
On the framework side, Next.js 16.4 shipped next upgrade --agent, which hands a coding agent version-specific upgrade guidance, codemods and verification steps, and an experimental agentUpgrade setting that nudges you during next dev and next build when a security-relevant upgrade exists - Next.js. Cache Components, now recommended for every Next.js app, will become the default in Next.js 17. For an app on the credit this matters mostly as maintenance: an agent can keep the framework current, which is one less reason a small project rots.
Gateways and hosting
Vercel's AI Gateway added Claude Haiku 5.5 on October 7 under the model string anthropic/claude-haiku-5.5, and says it passes through provider pricing with no markup and no platform fee on inference, including bring-your-own-key requests - Vercel. If you route Claude through the gateway with your own Anthropic key, the requests are sent with that key, so they should bill to the organization the key belongs to; make one test call and confirm on the Console's Cost page that the credit drew down before relying on it. The same day the gateway added OpenAI's Decisions API - Vercel. On October 9 Vercel let agents search, price and start domain purchases from its CLI while leaving the actual purchase confirmation to a human - Vercel.
Databases moved in the same direction. On October 2 Supabase announced it was acquiring Turso, whose SQLite-based platform loads databases on demand and suspends them when idle, so that agents can create databases "as easily as creating a file"; Supabase says it already launches more than one million databases a week - VKTR. For a credit-funded app the relevance is cost shape: per-user or per-agent databases that cost nothing while idle fit the same "pay only when used" logic as the API credit.
Getting paid by agents
At the other end of the transaction, Cloudflare's Birthday Week shipped a closed beta that lets sellers charge AI agents per request. The Monetization Gateway answers an unpaid request with HTTP 402 Payment Required inline, the agent signs a payment authorization, and settlement runs through Coinbase's x402 facilitator in USDC on Base; it is open to eligible U.S. sellers and buyers - Cloudflare. The same week Cloudflare released Vinext 1.0, a Vite-based way to run Next.js apps outside Vercel - Cloudflare. Both were part of a week of 46 announcements - Cloudflare Birthday Week recap.
This closes a loop for anyone building on the credit. The credit pays for your app's AI; per-request payment rails let agents pay for your app's output. A small API or data product could plausibly run its inference on the credit and charge agent buyers per call, with neither side holding a subscription. Our guide to selling to AI agents covers the payment protocols in depth, and our guide to shipping an MCP server for your product covers how agents discover and call what you build.
How the credit compares
Anthropic is not the only lab bundling developer money into consumer plans. Since January, Google has included $10 a month in Google Cloud credits with Google AI Pro and $100 a month with Google AI Ultra, usable for deploying to Vertex AI or Cloud Run and for Gemini API usage - Google. The structural difference is scope: Google's credit is general cloud money spent through a billing account, while Anthropic's is model money spent in a Console organization. For a builder already on Claude, Anthropic's version is the more direct fit, and on Max it returns the full plan price as credit.
One calendar date is worth noting for anyone using the credit to get a product in front of investors. Y Combinator's Winter 2027 deadline is November 2 at 8 p.m. PT, with decisions by December 11 - Y Combinator. That is one full credit cycle away: enough to ship the reference build from Section 8, put it in front of real users, and apply with usage numbers instead of a plan. Our analysis of what YC is funding in 2026 covers the categories that keep getting funded. Why this matters: the credit is most valuable when it turns into evidence. How to apply it: pick a date, ship by it, and let the month's usage tell you whether to keep paying once the credit is gone.
13. A Decision Framework
The credit rewards a clear plan and punishes drift, because it expires every cycle whether or not you used it. The right plan depends less on which plan you pay for than on what you are trying to build, so the framework below starts from your situation. In every case the first move is the same: claim the credit, because an unclaimed credit is simply money left on the table, and claiming it into a fresh organization costs nothing and commits you to nothing beyond the one-time choice of organization.
The second principle is to separate building from running. Your Max plan pays for building in Claude Code; the credit pays for what your product calls once it exists. Every decision below follows from keeping those two pools apart, putting the credit only where your own key sits in the runtime path, and capping every workspace so that no single mistake can spend more than a known amount.
If you only use Claude Code today, claim into a new organization anyway and give the credit one recurring job, such as a weekly batch digest of your notes, tickets or reading list on Haiku 5.5. It costs cents, it proves the setup works, and it means the next idea you have already has a funded, capped workspace waiting for it.
If you have an app idea, ship the Section 8 build: one Claude-powered feature on Haiku 5.5, a per-user limit in your own code, and a workspace with a spend limit. Use Claude Code on your plan to write it, and let the credit pay for what users trigger. Move a feature to Sonnet 5.5 only when you can show Haiku's answers are not good enough.
If you already run a paid product, link its existing organization so the credit offsets the bill automatically, then move anything that can wait into the Batches API and cache every stable prefix. For a product with real users this is the highest-value option, because every dollar of credit replaces a dollar you would have paid anyway.
If you run a Team plan, treat the pool as a shared sandbox rather than a per-person allowance, give each project its own capped workspace, and remember that a pool over $200 lifts the organization to at least the Build tier, which doubles its monthly spend cap from $500 to $1,000.
If you are building an agent, start on Managed Agents if you do not want to operate sandboxes and on the Agent SDK if you do, and cap tools, turns and workspace spend either way. Judge the agent by cost per finished task, not cost per request.
Across all five, the measurements that decide what happens after the first month are the same: cost per active user, cost per finished task, and whether anyone would pay for the result. If the credit covers your costs and users keep coming back, you have a product that is cheap to run; if the credit runs out because users love it, you have a pricing question rather than a cost problem, which is the better problem to have. If nobody uses it, the month cost you nothing but time, and the next cycle's credit is a free second attempt.
The broader lesson of October 7 is that the price of useful intelligence keeps falling while the tools to ship it keep getting simpler. A $100 credit on Haiku 5.5 covers a short-prompt workload that would have needed $1,000 on Haiku 4.5, and a framework, a database and a host can now be assembled by an agent in an afternoon. The scarce ingredient is no longer budget or tooling. It is a specific product that a specific group of people wants, and the credit is best spent finding out whether you have one.
This guide reflects Anthropic's credit terms, model prices and the related tools as of October 10, 2026. The credit is still rolling out, and pricing and terms change frequently, so check the Claude Help Center and the Claude Platform docs before relying on any figure here.