How OpenAI's new plan-usage permission works, what it costs you, what it will not do, and when to let your users' ChatGPT plan pay for your app's AI
On September 29, OpenAI turned a login button into a payment method.
At its DevDay keynote, OpenAI extended Sign in with ChatGPT so that a user can let a third-party app spend the AI usage already included in their ChatGPT plan. The user clicks Continue with ChatGPT, approves a permission called "Use your ChatGPT plan", and from then on the app's eligible requests to OpenAI's Responses API come out of the user's subscription instead of the developer's API bill. OpenAI launched it with 16 partners, including Cognition's Devin, Notion, Vercel, T3, OpenClaw and Dactyl - OpenAI DevDay 2026 recap. Sam Altman put the pitch to developers in one sentence on stage: "They're already paying for an AI subscription, and now you don't have to cover their token costs to get them going" - The New Stack.
But the button is not a free lunch. Commercial apps can only use the payment side through a limited trial, the plan route strips out sampling controls and most hosted tools, the user (not you) sets a weekly cap on how much of their plan your app may spend, and you tie a core part of your unit economics to one AI lab's policy. Anthropic, notably, forbids the equivalent move with Claude subscriptions entirely. Whether this is a gift or a trap depends on what your app does, who your users are, and how much of your margin inference eats today.
This guide covers exactly how the permission works under the hood, a working integration with code, the preview limits that decide which apps it suits, the money math of who pays for a user, the six other ways to let users pay for AI, the failure modes, and a decision framework you can apply this week. Everything here was verified against OpenAI's own documentation and the live pages cited inline on October 7, 2026.
Contents
- What OpenAI Shipped on September 29
- Who Pays for Inference: The Structural Question
- How Sign in with ChatGPT Works Under the Hood
- Build It: A Working Integration, Step by Step
- What the Plan Route Will Not Do
- Limits, Caps and Credits: The User Holds the Meter
- The Money Math: What One User Costs You
- The Other Ways to Let Users Pay, Compared
- Who Has Shipped It, and How They Price It
- Where It Fails: Risks, Abuse and Platform Dependence
- Agents Change the Math
- A Decision Framework for Your App
The Seven Ways to Pay for an App's AI, Scored
Before going deep on OpenAI's feature, it helps to see it next to every other way an app can pay for the AI it uses. The table below scores seven billing paths on the five things a founder actually trades off: how much inference cost stays on your books, how much effort it takes a user to start paying, how much of the model capability you keep, how much you depend on someone else's rules, and how many of your users can use the path today. Each cell holds the score and the fact behind it, and the details for every row follow in Sections 5 through 8.
| # | Billing path | What It Does | Cost exposure (25%) | User friction (25%) | Capability (25%) | Control (15%) | Reach (10%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | Developer pays (your API key) | You buy tokens, price them into your plans | 2 - you pay every token; 7.4% of AI sign-ups show multi-account abuse | 8 - user does nothing, but heavy use needs your paid plan | 10 - any model, every hosted tool and parameter | 8 - you pick and switch providers | 10 - every user | 7.2 |
| 2 | On-device model (Apple) | Apple's framework runs a local model at no inference cost | 10 - "AI inference that is free of cost" on device | 10 - nothing for the user to do | 3 - a 3B-parameter local model, Apple devices only | 5 - Apple's platform and rules | 4 - Apple Intelligence devices only | 6.9 |
| 3 | Sign in with ChatGPT | User's ChatGPT plan pays for eligible Responses API calls | 9 - inference leaves your bill; hosting stays | 9 - one click, no key, no card | 5 - OpenAI models only, no hosted tools or temperature | 3 - commercial use behind a limited trial | 5 - paid ChatGPT subscribers; open-source tools wider | 6.7 |
| 4 | OpenRouter OAuth | User connects an OpenRouter account and pays from its credits | 9 - user's credits; 5.5% fee on their purchases | 4 - needs an OpenRouter account with credits | 8 - many labs' models through one API | 7 - open OAuth, no approval gate | 2 - mostly developers have accounts | 6.5 |
| 5 | Classic BYOK API key | User pastes their own provider API key | 10 - provider bills the user directly | 2 - create an API account, add billing, paste a key | 9 - the full provider API | 6 - you hold secrets that leak | 2 - technical users only | 6.4 |
| 6 | Puter.js user-pays | User's Puter account covers AI calls from your frontend | 9 - "usage never touches your bill" | 6 - one sign-in, free monthly allowance first | 6 - Puter's menu, client-side only | 5 - depends on a small vendor | 3 - an unfamiliar brand to most users | 6.3 |
| 7 | Hugging Face OAuth | inference-api scope bills the signed-in HF user | 9 - user billed at provider rates, no HF markup | 4 - HF account and credits needed | 7 - open-weight models via Inference Providers | 7 - standard OAuth, no gate | 2 - mostly ML practitioners | 6.3 |
How to read the criteria. Cost exposure (25%) asks how much of the inference bill stays with you; a 10 means none. User friction (25%) asks how many steps stand between a new user and paid-for AI. Capability (25%) asks how much model quality, model choice and API surface you keep. Control (15%) asks how exposed you are to another company's approval, pricing or policy changes. Reach (10%) asks what share of a typical app's users can use the path today. Puter (6.30) edges Hugging Face (6.25), so both display as 6.3.
The headline from the table is that Sign in with ChatGPT is a complement, not a replacement for paying for your own AI. Developer-paid inference still wins overall because nothing else gives you every model and every feature for every user, and on-device inference wins on cost wherever a small local model is good enough. What OpenAI's feature uniquely offers is the combination of near-zero user friction and near-zero cost exposure for one specific, valuable audience: people who already pay OpenAI every month. Most apps should treat it as a second billing path, not the only one.
1. What OpenAI Shipped on September 29
To understand what changed, separate the two things that share one name. Sign in with ChatGPT as an identity provider is not new: it began rolling out in beta on July 29, 2026 across select plugins and partner sites, starting with Airtable, GitLab, HubSpot, Notion, Supabase and Vercel - ChatGPT changelog. That version works like "Sign in with Google": the app receives a name, an email address and a profile picture, and nothing else. What DevDay added is a second, optional permission on top of the login that lets the app spend the user's plan.
OpenAI's recap describes the new part plainly: users can spend their allowance across 16 partners, control how much each can use, and "eligible OpenAI usage counts toward your plan limits." The same recap states that identity is available globally while plan usage is available to Plus and Pro users in participating tools - OpenAI DevDay 2026 recap. OpenAI's Help Center, by contrast, lists Go, Plus and Pro as the plans that can use their plan in participating commercial apps, and adds that supported open-source tools remain available to all ChatGPT users - OpenAI Help Center. The developer docs still say Plus and Pro, so check eligibility against the Help Center before you promise Go users anything.
The idea itself has been in motion for more than a year. In May 2025, OpenAI published an interest form for developers who wanted to let users sign in to third-party apps with ChatGPT, asking about app sizes from fewer than 1,000 to more than 100 million weekly users, shortly after previewing a ChatGPT sign-in inside Codex CLI - TechCrunch. By August 2025, Codex's IDE extension and CLI offered "Sign in with ChatGPT" as one-click authentication that removes API keys - ChatGPT changelog. The coding tools came first because their users already lived inside ChatGPT plans; DevDay generalized the pattern to other companies' apps.
The image below is OpenAI's own illustration of the new consent step, from the DevDay recap. The toggle labelled "Use your ChatGPT plan" is the whole feature: without it, the user has only signed in.
Notice the wording under the toggle: the app may "consume usage from your ChatGPT Plan's limits". That phrase carries the economics of the whole feature. The app is not charging the user's card and it is not receiving money; it is drawing down an allowance the user already bought, inside limits the user controls. Notice too the line at the bottom, which tells users they can manage login connections in ChatGPT settings. The control surface lives at OpenAI, not in your app.
The launch partners
OpenAI's public directory sorts participating apps into three groups. Apps with ChatGPT plan usage: Amp Code, Conductor, Dactyl, Devin, Hermes Agent, Hyperagent, Kilo Code, Lovable (listed as coming soon), Notion, Vercel, Vorflux and Warp. Apps with sign-in only: Airtable, Canva, GitLab, HubSpot and Supabase. Open-source integrations: OpenClaw, OpenCode, Pi and T3 - ChatGPT Learn directory. Coding agents dominate the first group, and Section 5 explains why that is no accident.
The access model for developers is lopsided by design. Sign in with ChatGPT for websites and plugins "is currently available to selected commercial partners through a limited trial", while ChatGPT plan usage "is available to all open-source partners and selected private clients" - OpenAI Quickstart. Anyone shipping a paid or remotely hosted app is pointed to an interest form - OpenAI client ID request. So the precise reading of "available" on October 7 is: open-source and local tools can ship it now, commercial apps can apply.
The same week reset the price of mid-tier AI
The feature landed in a week that also repriced the models it would carry. OpenAI released GPT-6.1 Sol the same day at $2 per million input tokens, $0.10 cached and $10 per million output tokens for prompts up to 272K tokens - OpenAI API changelog. A day earlier, Anthropic had priced Claude Sonnet 5.5 at the same $2 and $10 - Anthropic. A day later, Google announced Gemini 4 Argon at an introductory $2 and $10, though only a set of trusted cyber defenders in its Fairwind program can use it so far - Google.
That convergence matters for this guide because it sets the yardstick. When three labs price their strong mid-tier models identically, the cost of serving a user stops depending much on which lab you pick and starts depending on how much the user does and who pays for it. OpenAI's DevDay answer to the second question is the plan route. The keynote below covers the full set of announcements, including GPT-6.1 Sol and the plan usage partners.
The keynote frames Sign in with ChatGPT as one piece of a larger move to open ChatGPT as a shared surface for developers, alongside always-on agents called dots and new plan tiers. For a builder, the useful takeaway is narrower: OpenAI now wants its subscribers' allowance to travel with them into other products, and it has published the exact contract under which your app can receive it.
2. Who Pays for Inference: The Structural Question
Strip away the branding and every AI app answers one structural question: who pays for the inference? Each request to a model consumes compute that someone must buy. For most of the short history of AI apps, the answer was the developer. You held one API key, every user's requests ran through it, and you priced your product to cover the bill plus a margin. That model is simple and it gives you full control, but it makes your gross margin a direct function of your users' behavior, which no traditional software company ever had to live with.
The numbers show how heavy that weight is. ICONIQ's 2026 State of AI report, surveying about 300 executives, found AI product gross margins at 45% in 2025, projected to reach 53% in 2026 and 59% in 2027 - ICONIQ. Bessemer's 2025 analysis put the average gross margin of its fastest-growing "Supernova" AI startups at roughly 25%, often negative - Bessemer Venture Partners. Classic SaaS businesses are used to margins far above either figure. The difference is inference, and it grows with your best customers rather than shrinking.
The trend is upward, helped by cheaper models and better routing, but even the 2027 projection sits well below what software investors historically expected. We covered the full set of pricing levers (credits, tiers, outcome pricing and the margin floor that stops one power user from sinking you) in our guide to pricing an AI product to beat token costs. Sign in with ChatGPT attacks the same problem from a different side: rather than pricing around the cost, it moves the cost off your books for the users who opt in.
Why OpenAI wants to carry your bill
Ask why the lab would do this, and the answer comes from the shape of a subscription. A ChatGPT plan is a flat monthly fee for a bounded allowance. Some subscribers use all of it; many use a fraction. Every unit of allowance that goes unused is pure margin for OpenAI, but it is also value the subscriber did not feel, and subscribers who feel little value cancel. Letting the allowance flow into other apps raises the felt value of the same $20 without raising OpenAI's price, which is the strongest retention tool a subscription business has.
There are two further prizes. The first is identity: "Sign in with Google" and "Sign in with Apple" made those companies the front door to millions of apps, and each login is a daily reminder of whose account you live in. OpenAI's recap counts 1.2 billion weekly users across its products - OpenAI DevDay 2026 recap, and a login button lets it extend that relationship into every partner app. The second is demand capture: every request your users make on their plan runs on an OpenAI model, which makes it harder for you to route that user to a competitor's model later.
Why developers want it anyway
From the developer's side, the logic runs the other way and is just as strong. Your cheapest possible customer acquisition is a user who can start using your product's AI without a credit card, without an API key, and without you underwriting the cost of a free trial. Free trials for AI products are expensive precisely because inference is real money, and they are abused precisely because free tokens are worth stealing (Section 10 has Stripe's numbers). A user who pays with an existing ChatGPT plan has already been identified, already passed OpenAI's payment checks, and already accepted OpenAI's usage policies.
The diagram shows the essential rewiring. In the old model, the inference cost sits between your revenue and your margin. In the new one, inference is settled between the user and OpenAI, and your app is paid only for what it uniquely adds. That is a cleaner business in theory: your price reflects your product, not a resold commodity. In practice it means your product must stand on its own value, because "we include the AI" stops being a reason to pay you.
The analogy, tested
The closest historical analogy is the shift in mobile from apps that bundled data costs to apps that simply used the user's own data plan. Nobody expects a maps app to pay for the user's mobile data; the carrier sells the bandwidth and the app sells the experience. AI inference is moving toward the same split, where the lab sells the intelligence and the app sells the outcome. The analogy holds on the economics: inference, like bandwidth, is a metered utility that the end user increasingly buys in bulk.
It breaks in one important place. Mobile data is a neutral pipe: any app works on any carrier. A ChatGPT plan is not neutral: it only pays for OpenAI models, under OpenAI's rules, in OpenAI's format. If the analogy were exact, a user's AI plan would work with any model in any app. Today it does not, and the gap between the utility the analogy promises and the walled allowance that exists is exactly where the risks in Section 10 live. Keep that gap in mind as we go through the mechanics.
3. How Sign in with ChatGPT Works Under the Hood
Underneath the friendly button, Sign in with ChatGPT is standard web plumbing: OAuth 2.0 with PKCE for authorization and OpenID Connect for identity. If you have ever added "Sign in with Google" to an app, roughly 80% of this will feel familiar. The new 20% is a second set of permissions that turns a login into a payment channel, and that new part is where every interesting design decision lives. Understanding the plumbing matters because the failure modes (expired tokens, revoked consent, a user who hit their limit halfway through a streamed answer) all show up as specific technical events your app has to handle.
The first thing to internalize is that OpenAI split the feature into two separate grants. The identity grant uses the scopes openid profile email and gives your app a stable account identifier, the person's name, email address and profile picture, and nothing else - OpenAI Quickstart. The plan-usage grant adds three more scopes, offline_access resource.invoke chatgpt.tokens.use.direct, bound to the resource https://api.openai.com/v1, and only when the user approves those does the token response include an access token that can spend their plan - OpenAI registration guide. A user can sign in and decline plan usage, and your app must treat those as two different states rather than one.
Three integration surfaces, three different doors
OpenAI documents three ways in, and they are not equally open. A website integration is a classic confidential or public OIDC client: you request a client ID (typically beginning with oaiapp_), register exact callback URLs, and your backend exchanges the code and verifies the ID token - OpenAI website guide. A ChatGPT plugin integration nests two OAuth transactions inside each other: ChatGPT authorizes your connector (the outer flow), and your app signs the user in with ChatGPT (the inner flow), each with its own state, PKCE values and codes - OpenAI plugin guide. Both of those currently require OpenAI to approve you as a partner.
The third door, the open-source client flow, is the one anybody can walk through today. It uses dynamic client registration: your tool starts its very first authorization with client_id=dynamic_agent_client, the user names and approves the agent, and the callback hands back a freshly issued client ID bound to that user and workspace. No client secret and no partner API key are needed - OpenAI registration guide. The catch is in the fine print of the overview: these docs cover open-source and locally hosted apps, and a paid or remotely hosted app is pointed at an interest form instead - OpenAI plan usage overview.
- Website flow gives identity to commercial partners in the limited trial
- Plugin flow gives one-click connector auth inside ChatGPT
- Open-source flow gives plan usage to any open-source or local tool
This asymmetry is the most important practical fact in the whole guide. If you ship a command-line tool, a desktop app or a self-hosted agent under an open-source license, you can implement plan usage this week. If you run a closed-source SaaS on your own servers, you can build against the same contract but you will need OpenAI's approval before real users can pay with their plan. Plan your roadmap around that gate, not around the demo you saw at DevDay.
Whichever door you use, the button itself is meant to look the same everywhere. OpenAI's Help Center illustrates the expected placement with a fictional app: the ChatGPT option sits beside the app's own email sign-in with comparable prominence, labelled Continue with ChatGPT, which is also the label OpenAI's guidelines ask partners to use.
The avatar on the right of the button shows the ChatGPT account the browser is already signed in to, which is the friction advantage in miniature: for a logged-in ChatGPT user, signing up for your app is one click. What the image cannot show is the second screen, where the user decides whether your app may also spend their plan.
Clients versus hosts
The open-source flow introduces a concept most OAuth integrations never need: the difference between a client and an agent host. The client is the OAuth registration the user authorized, and its issued client_id carries that user's plan-usage settings and limits. A host is a place where an instance of your tool runs, such as a laptop install or a self-hosted VM, and each host sends its own stable, opaque ext_agent_host_id so usage can be attributed per machine while the limits stay shared - OpenAI plan usage overview.
OpenAI accepts three host-ID formats. The recommended one is a JWK thumbprint URI derived from a public key as described in RFC 9278, persisted as urn:ietf:params:oauth:jwk-thumbprint:...; a urn:uuid: value generated once per host and a did:key identifier are also supported. OpenAI is explicit that a key-derived host ID is currently only an identifier: it does not verify that the host holds the private key. Treat the host ID as a label for attribution, never as a credential, and never put an email address or user ID into it.
The diagram shows the happy path and the one unhappy path you will hit most often. Notice that billing never appears as a separate step: there is no checkout and no card, because the user's consent screen is the purchase decision. That is precisely why the error branch matters so much. When the plan runs out, your app receives a usage-limit error rather than a bill, and OpenAI states plainly that it does not silently switch the request to another billing path - OpenAI errors guide.
Tokens, lifetimes and what is inside them
Access tokens last one hour (expires_in: 3600). Refresh tokens last 30 days, and every successful refresh returns a replacement refresh token with a fresh 30-day lifetime, with no fixed cap on how many times that can repeat - OpenAI token reference. In practice that means a user who opens your tool at least once a month stays connected indefinitely, while one who disappears for five weeks has to sign in again. The access token is a JWT whose audience is https://api.openai.com/v1, whose issuer is https://auth.openai.com, and which carries an opaque, encrypted block of OpenAI authentication metadata that your app is told to treat as a black box.
Refresh tokens rotate, which creates a classic race condition: if two processes refresh the same session at the same moment, one of them presents a token that has already been replaced. OpenAI's guidance is to serialize refreshes per session and always store the newest replacement - OpenAI accounts and sessions guide. If your app runs several worker processes, put a lock around refresh from day one; a refresh_token_reused error in production is almost always this bug.
4. Build It: A Working Integration, Step by Step
This section walks through the open-source client flow end to end, because it is the path any developer can ship without waiting for approval, and because a commercial integration reuses most of the same pieces. The examples are in Python with the widely used requests and PyJWT libraries, and each step maps to a specific page of OpenAI's documentation. The goal is not to replace those docs but to show the shape of the code and, more usefully, where the sharp edges are.
Before writing any code, decide two product questions, because they change the code. First, what does your app do when a user signs in but declines plan usage: offer their own API key, offer your paid credits, or show a read-only mode? OpenAI expects a clear choice here, noting that "a coding harness still needs a way to pay for inference" - OpenAI errors guide. Second, how many machines will one user run your tool on? That decides how you generate and store host IDs.
Step 1: Create a stable host ID
Generate the host identifier once per installation and persist it before the first sign-in. A UUIDv4 with a urn:uuid: prefix is the simplest accepted format:
import json, os, uuid, pathlib
CONFIG_DIR = pathlib.Path.home() / ".config" / "mytool"
HOST_FILE = CONFIG_DIR / "host.json"
def get_host_id() -> str:
CONFIG_DIR.mkdir(parents=True, exist_ok=True)
if HOST_FILE.exists():
return json.loads(HOST_FILE.read_text()) ["ext_agent_host_id"]
host_id = f"urn:uuid:{uuid.uuid4()}"
HOST_FILE.write_text(json.dumps({"ext_agent_host_id": host_id}))
os.chmod(HOST_FILE, 0o600)
return host_id
The important property is stability, not secrecy. Restarting your tool or signing in again must not mint a new host, because OpenAI uses the host ID to attribute usage, and a tool that creates a fresh host on every launch will look like hundreds of machines to the user reviewing their usage. If you later want the stronger format, derive a JWK thumbprint from a key pair you already keep for the host and persist that URI instead.
Step 2: Start authorization
Generate a fresh state, OIDC nonce and PKCE verifier for every attempt, start a loopback listener, then open the system browser at the authorize endpoint. First-time registration uses the special client ID and your app's name as a hint; later sign-ins use the issued client ID and omit the hint - OpenAI registration guide.
import base64, hashlib, secrets, urllib.parse, webbrowser
AUTHORIZE = "https://auth.openai.com/api/accounts/authorize"
REDIRECT = "http://127.0.0.1:1455/auth/callback"
SCOPES = "openid profile email offline_access resource.invoke chatgpt.tokens.use.direct"
def start_auth(host_id: str, issued_client_id: str | None = None) -> dict:
verifier = secrets.token_urlsafe(64)
challenge = base64.urlsafe_b64encode(
hashlib.sha256(verifier.encode()).digest()
).rstrip(b"=").decode()
attempt = {"state": secrets.token_urlsafe(24),
"nonce": secrets.token_urlsafe(24),
"verifier": verifier}
params = {
"client_id": issued_client_id or "dynamic_agent_client",
"ext_agent_host_id": host_id,
"response_type": "code",
"redirect_uri": REDIRECT,
"scope": SCOPES,
"resource": "https://api.openai.com/v1",
"state": attempt ["state"],
"nonce": attempt ["nonce"],
"code_challenge": challenge,
"code_challenge_method": "S256",
}
if issued_client_id is None:
params ["agent_name_hint"] = "MyTool" # first registration only
webbrowser.open(AUTHORIZE + "?" + urllib.parse.urlencode(params))
return attempt
Two details in the redirect URI trip people up. The callback must use 127.0.0.1, never localhost, and only the port may change between sign-ins: /auth/callback must stay exactly /auth/callback. Start the listener before opening the browser, or a fast user will land on a dead page.
Step 3: Handle the callback and exchange the code
When the browser returns, validate state first, handle error=access_denied without exchanging anything, and keep the issued client_id from a new registration (it looks like oaiapp_...). Then post a form-encoded authorization_code grant to the token endpoint, with no client secret:
import requests
TOKEN = "https://auth.openai.com/api/accounts/oauth/token"
def exchange(code: str, client_id: str, attempt: dict) -> dict:
resp = requests.post(TOKEN, data={
"grant_type": "authorization_code",
"client_id": client_id, # the ISSUED id, never dynamic_agent_client
"code": code,
"code_verifier": attempt ["verifier"],
"redirect_uri": REDIRECT,
"resource": "https://api.openai.com/v1",
}, timeout=30)
resp.raise_for_status()
return resp.json() # access_token, refresh_token, id_token, scope, expires_in
If the exchange fails with invalid_grant, discard that code and start a fresh authorization; codes are single-use. On a reauthorization the callback may omit client_id, so hold on to the one you sent, and reject the result if the callback returns a different ID rather than silently replacing the registration.
Step 4: Validate the ID token and the plan permission
Verify the ID token's signature against OpenAI's published JWKS, and check issuer, audience (your issued client ID), expiry and the nonce you generated. Then, separately, check that the granted scopes include chatgpt.tokens.use.direct, because a perfectly valid ID token proves only who the user is, not that they agreed to let you spend their plan:
import jwt # PyJWT
JWKS = jwt.PyJWKClient("https://auth.openai.com/.well-known/jwks.json")
def validate(tokens: dict, client_id: str, nonce: str) -> dict:
key = JWKS.get_signing_key_from_jwt(tokens ["id_token"])
claims = jwt.decode(tokens ["id_token"], key.key, algorithms= ["RS256"],
audience=client_id, issuer="https://auth.openai.com",
leeway=5)
if claims.get("nonce") != nonce:
raise ValueError("nonce mismatch")
granted = set(tokens.get("scope", "").split())
plan_enabled = "chatgpt.tokens.use.direct" in granted
return {"sub": claims ["sub"], "email": claims.get("email"),
"plan_enabled": plan_enabled}
The algorithm list is an assumption worth checking against the key metadata in the JWKS your app actually receives; pin whatever OpenAI publishes rather than accepting any algorithm. The plan_enabled flag is the branch point for your whole billing logic. When it is false, keep the user signed in and route them to your alternative payment path; when it is true, you can make inference calls on their plan.
Step 5: Store credentials safely and refresh them
OpenAI's documented storage pattern is one protected credential record per issued client ID, written atomically with owner-only 0600 permissions under an app-defined path such as ~/.config/YOUR_TOOL/, never committed and never logged - OpenAI registration guide. Refresh near expiry with a refresh_token grant that sends the issued client ID, the saved refresh token and the same resource, omitting scope so the original grant is retained:
def refresh(record: dict) -> dict:
resp = requests.post(TOKEN, data={
"grant_type": "refresh_token",
"client_id": record ["client_id"],
"refresh_token": record ["refresh_token"],
"resource": "https://api.openai.com/v1",
}, timeout=30)
resp.raise_for_status()
fresh = resp.json()
record.update(access_token=fresh ["access_token"],
refresh_token=fresh ["refresh_token"], # rotates every time
expires_in=fresh ["expires_in"],
scope=fresh.get("scope", record.get("scope")))
return record # write the whole record back atomically
On sign-out, revoke the renewable session at the revocation_endpoint listed in https://auth.openai.com/.well-known/openid-configuration, then clear the tokens locally. An empty HTTP 200 counts as success, even for a token that was already invalid. If revocation cannot be confirmed because of a network failure, OpenAI's guidance is to clear tokens anyway and tell the user they can disconnect the app in ChatGPT settings - OpenAI accounts and sessions guide.
Step 6: List models and make the first inference call
Populate your model picker from the signed-in account's own catalog rather than a hard-coded list, because what a Go user can reach differs from what a Pro user can reach. The documented call is a plain GET /v1/models with the user's access token, keeping models whose visibility is "list" - OpenAI models and inference guide. Inference then goes to the public Responses API, with two non-negotiable flags:
curl --no-buffer https://api.openai.com/v1/responses \
-H "Authorization: Bearer ${ACCESS_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6.1-sol",
"input": [{"role": "user", "content": "Summarize this pull request."}],
"store": false,
"stream": true
}'
The official Python SDK works unchanged if you pass the OAuth access token as api_key, since the SDK simply sends it as a bearer credential. Every request in this flow must set store: false and stream: true, and success means one thing only: receiving the response.completed event. A usage-limit failure can arrive as response.failed after tokens have already streamed, so a partially rendered answer is not a finished one, and your UI needs a state for "stopped because the plan ran out".
Step 7: Map every error to a user-visible state
The plan route returns a family of specific error codes, and each has a different correct response. Treating them all as a generic retry is the most common way to annoy users, because some of them can never succeed on retry:
| Code | HTTP | What your app should do |
|---|---|---|
subscription_sharing_user_not_eligible | 403 | Explain the restriction; do not loop through OAuth |
subscription_sharing_usage_limit_exceeded | 429 | Pause plan requests, link to ChatGPT usage settings |
subscription_sharing_usage_unavailable | 503 | Keep credentials, retry later with backoff |
subscription_sharing_unsupported_capability | 400 | Remove the parameter named in error.param |
subscription_sharing_invalid_user | 401 | Ask the user to sign in again |
The table condenses OpenAI's own recovery guidance - OpenAI errors guide. Two entries deserve emphasis. The 429 does not mean the user's whole plan is empty: an app-specific weekly cap can trigger it while the user still has ChatGPT usage left, so never tell the user "your ChatGPT plan is exhausted" on the strength of this code alone. And before a stream opens, admission failures may come back as a bare {"detail": "..."} body rather than the standard error object, so parse defensively. OpenAI also does not notify your app when a user disconnects it in ChatGPT settings; you only learn that when the next request or refresh fails.
If your product is built around an agent harness rather than direct API calls, OpenAI documents one more shortcut: Codex app-server can be pointed at the Responses API with the user's OAuth token supplied through an environment variable, so a tool that embeds Codex can bill the user's plan without a separate Codex sign-in - OpenAI Codex app-server guide. The token renewal remains your app's job, and you restart the app-server process with the new token and resume the saved thread.
5. What the Plan Route Will Not Do
Here is the part of the launch that the keynote did not dwell on. A request billed to a user's ChatGPT plan is not the same product as a request billed to your API key, even when both hit the same endpoint with the same model name. OpenAI labels the current behavior a preview and lists its limits plainly - OpenAI preview limitations. If your app's value depends on any of the missing pieces, you need a second billing path for those features, and that changes your architecture.
The restrictions fall into four groups. Several request fields are simply rejected: background, conversation, max_output_tokens, max_tool_calls, metadata, moderation, multi_agent, prompt, prompt_cache_retention, safety_identifier, temperature, top_logprobs, top_p, truncation and user. Conversation state cannot live on OpenAI's side over HTTP, because previous_response_id is off the table and every request must carry its own history in input. Explicit system-role message items are rejected, so instructions go in the instructions field or a developer message. And the hosted tools that make the Responses API powerful are mostly unavailable.
| Capability | Your API key | User's ChatGPT plan |
|---|---|---|
| Function and custom tools | Yes | Yes (namespaced) |
| Web search | Yes | Subject to account policy |
| File search, Code Interpreter | Yes | No |
| Image generation | Yes | No |
| Hosted MCP, connectors, tool search | Yes | No |
| Native computer use | Yes | No |
| Audio or video input, Files API | Yes | No |
temperature, max_output_tokens | Yes | No |
| Server-side conversation state | Yes | No (store: false) |
The contrast is sharp because the model itself is capable: GPT-6.1 Sol's own model page lists web search, file search, image generation, Code Interpreter, hosted shell, computer use, MCP and tool search among its supported Responses tools - OpenAI GPT-6.1 Sol model page. Through the plan route, most of that list disappears.
What does this mean in practice? The plan route is shaped for one kind of application: a client-side agent that runs its own tools. A coding agent that reads files, runs shell commands and calls local MCP servers through function calls loses almost nothing, which is exactly why the first wave of partners is dominated by coding tools. A chat product that relies on OpenAI-hosted retrieval, image generation or sandboxed code execution loses its differentiating features on the plan route. And a product that needs deterministic output via temperature: 0, or a hard ceiling on output length via max_output_tokens, has to enforce those properties another way (constrained prompts, structured outputs, client-side truncation) or keep those calls on its own key.
The missing safety_identifier and user fields deserve a separate note for anyone who runs a consumer product. On your own key, those fields help OpenAI attribute abusive requests to a specific end user instead of to your whole organization. On the plan route, the request already belongs to an identified ChatGPT account, so the attribution problem moves to OpenAI. That is a quiet but real shift of trust-and-safety burden, and it is one reason the arrangement can be attractive for apps with a free tier, a point Section 10 returns to.
6. Limits, Caps and Credits: The User Holds the Meter
The single biggest conceptual shift in Sign in with ChatGPT is that the meter moves to the user. On your own API key, you decide how much a customer can consume and you see every token on your bill. On the plan route, the user decides, inside ChatGPT, how much of their plan your app may consume, and your app learns about that decision only when a request is refused. Building a good product on top of that requires understanding exactly what the user can control.
The user-facing controls live in ChatGPT under Settings, then Usage. Each connected app gets a weekly usage limit expressed as a percentage of the user's overall weekly ChatGPT usage, adjustable from 10% to 100% - OpenAI accounts and sessions guide. OpenAI stresses that this limit is a cap, not a separate pool: it does not reserve usage for the app or add anything to the plan, and an app can hit its own cap while the user still has ChatGPT usage left - OpenAI Help Center. App requests count against the same allowance the user spends on Codex and ChatGPT Work, so a heavy day in your app is a lighter day in Codex.
The five-hour window and why Plus users will feel it
The plan tiers do not experience this equally. For ChatGPT Plus users, a five-hour usage limit is shared across every app where they use their plan, including private and open-source clients, so no app gets its own allowance inside that window; Pro users are exempt from the five-hour limit, though weekly limits still apply. OpenAI's own usage estimates make the scale concrete: on Plus, a five-hour window covers roughly 15 to 160 local messages with GPT-6.1 Sol, 5 to 45 with GPT-6 Astra, and 350 to 3,000 with GPT-6 Luna, with the spread depending on the task - OpenAI Codex pricing and limits.
Those ranges tell you which apps the plan route actually suits. A writing assistant that sends a few dozen modest requests a day will rarely touch the ceiling on Plus. An autonomous coding agent that fires a long chain of tool calls per task can burn a Plus window in a single session, and when it does, the user's other plan-powered tools stop too, because they share the same window. If your app is agent-heavy, design for the reality that many of your Plus users will run out mid-task, and make the experience of running out graceful rather than broken.
- Weekly app cap set by the user, 10% to 100%
- Five-hour window shared across all apps on Plus
- Weekly plan limit applies on every tier, Pro included
- Credits used only after an explicit opt-in
Read together, those four controls mean your app has no guaranteed budget at all. It has whatever the user chose to grant, minus whatever their other tools already spent this window. That is a real product constraint, but it is also an honest one: the user sees exactly where their allowance went, and an app that delivers value per unit of usage will look good in that view, while an app that wastes tokens on bloated prompts will look expensive in a place the user checks.
Credits, auto-purchase and the silent-charge trap
When a user exhausts their plan, an app can continue only if the user has explicitly allowed apps to draw on ChatGPT credits. That permission is off by default, applies across all participating apps, and only works for an app whose weekly limit is set to 100%; a lower cap blocks credit use for that app entirely - OpenAI Help Center. Credit consumption follows a published rate card: GPT-6.1 Sol costs 50 credits per million input tokens, 2.5 per million cached input tokens and 250 credits per million output tokens - OpenAI Codex pricing and limits. OpenAI does not publish one universal dollar price per credit, saying purchase prices and discounts depend on the plan or agreement.
There is one trap worth warning your own users about. The Help Center notes that if a user has automatic credit purchases enabled and allows apps to use credits, continued usage can trigger automatic charges without a separate notice when included usage runs out. A related misunderstanding drew press attention: signing out of a third-party app ends the session there but does not disconnect it from ChatGPT or stop its authorized use of the plan, which Notebookcheck highlighted in its coverage of the feature - Notebookcheck. Disconnecting happens in ChatGPT under Settings, then Security and login. If you want your users to trust the button, put both of those facts in your own help docs.
The UI OpenAI expects you to build
OpenAI publishes UI guidelines that read less like suggestions and more like the shape of an eventual review checklist. After the first sign-in with plan usage enabled, show a one-time confirmation along the lines of "You're using your ChatGPT plan" with a "Got it" action. When requests use the plan, display "Using ChatGPT plan" near the composer or model selector, with a "Manage usage" link to ChatGPT's usage settings. When a limit is hit, make "Manage usage" the primary action and your own credits, if you sell any, the secondary one - OpenAI UI/UX guidelines.
One guideline has direct commercial consequences: your app must clearly show which of its own plans support ChatGPT plan usage, on its pricing or plan-comparison page, with "Use your ChatGPT plan" in the feature list of each supported plan. OpenAI's illustrative example puts the feature on a $20 Pro tier and leaves it off the free tier. That example is explicitly non-binding, but it points at the pricing pattern most partners are likely to converge on, and Section 7 explains why the math pushes in that direction.
7. The Money Math: What One User Costs You
Every pricing decision about Sign in with ChatGPT comes down to one comparison: what a given user costs you in inference on your own API key, versus what they already pay OpenAI for a plan that could cover that cost. If the first number is small, the plan route mostly buys you a frictionless sign-up rather than savings. If the first number is large, the plan route can rescue your margins, but it also collides with the limits OpenAI places on each plan. Both halves of the comparison are now public, so the math can be done precisely.
Start with what the user already pays. ChatGPT's consumer plans are Go at $8 a month, Plus at $20 and Pro from $100 - OpenAI pricing. Pro now comes in three tiers at $100, $200 and $500, with the new $500 tier the only one that includes Astra Ultrafast, and new Pro 200 subscriptions receive a lower allowance than before unless they qualify for grandfathering until October 29, 2026 - OpenAI Help Center. For business buyers, a Business Premium seat costs $125 a month billed monthly and includes five times the usage of a Standard seat with no five-hour limit - OpenAI Help Center.
The spread is the first insight. A user on Pro 500 has prepaid a large AI budget and is likely a heavy, professional user; a Plus user has prepaid a modest one. Because plan usage draws from the same pool the user spends on Codex and ChatGPT Work, your app competes for that budget with OpenAI's own tools. An app that delivers more per unit of allowance than the user's alternatives will win a growing share of it; an app that burns allowance on long, repetitive context will get its weekly cap turned down.
What a user costs you on your own key
Now the other half. The worked example below prices four usage profiles at GPT-6.1 Sol's published list prices: $2 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache-write tokens and $10 per million output tokens - OpenAI GPT-6.1 Sol model page. The profiles are illustrative assumptions chosen to bracket common apps, not measured averages, so substitute your own telemetry before deciding anything.
The assumptions behind each bar are simple enough to check by hand. The light assistant makes 200 requests a month at 3,000 input and 600 output tokens each, which is 0.6 million input and 0.12 million output tokens, or $2.40. The daily writing app makes 1,000 requests at 8,000 input and 1,000 output tokens with no caching, for $26. The moderate agent runs 40 coding tasks a month, each reading 2 million input tokens across its turns (80% served from cache, the rest written to cache) and writing 50,000 output tokens, for $66.40. The heavy agent runs 150 such tasks, for $249. No single prompt in these examples crosses the 272K-token threshold where GPT-6.1 Sol's prices double.
Read against the plan prices, the chart splits users into two very different economic classes. For the light assistant, the user's inference costs you about one-eighth of a Plus subscription: the plan route saves you little money, and its real value is that the user never meets a paywall or a card form. For the agent profiles, the inference costs you more than the user's entire Plus plan costs them, which is precisely why a Plus user cannot sustain that workload on the plan route: OpenAI's estimates allow roughly 15 to 160 GPT-6.1 Sol messages per five-hour window on Plus. Heavy agent users belong on Pro tiers, and if your app attracts them, the plan route lets OpenAI's flat-fee subscription absorb what would otherwise be your largest variable cost.
A margin formula you can use
The cleanest way to reason about the decision is to split a user's cost into the part the plan route can move and the part it cannot. Suppose your app charges $20 a month, spends $8 a month per user on inference and $2 on hosting and other services. Your gross margin is ($20 minus $10) divided by $20, or 50%. If half of your inference spend comes from users who opt in to their ChatGPT plan, your inference cost falls to $4, and the same price now yields a 70% gross margin. Those numbers are illustrative, but the structure is general: the plan route raises your margin in direct proportion to the share of your inference that opted-in users generate.
That share is not the same as the share of users who opt in. Heavy users generate a disproportionate share of inference, and heavy users are also the likeliest to already pay for ChatGPT Plus or Pro, because they use AI all day. So a modest opt-in rate among your power users can move a large slice of your bill. Measure inference spend per user, rank users by it, and ask what fraction of your top decile already pays OpenAI. That single number tells you more about the value of this feature than any keynote.
Pricing your own plans around it
OpenAI's UI guidelines require you to show on your pricing page which of your own plans support ChatGPT plan usage. That rule effectively forces a pricing decision. Three patterns are emerging. The first is the bring-your-plan tier: a paid plan of yours that includes your product's features but relies on the user's ChatGPT plan for model usage, which lets you price it below an all-inclusive tier. The second is plan usage as a paid-tier feature: Dactyl's $20-a-month Builder plan, for example, runs its agent on the user's ChatGPT subscription and promises no token bills from Dactyl, while compiling, asset generation and publishing still consume Dactyl credits - The New Stack. The third is top-up credits: your own credits as the fallback when the user's plan runs out, which OpenAI's guidelines explicitly allow as a secondary action.
ICONIQ's survey suggests founders are already comfortable mixing models: companies blend an average of 1.7 pricing models, and consumption-based pricing rose from 35% to 42% in six months - ICONIQ. A plan-usage path slots naturally into that blend. If you meter usage today, our guide to setting up metered billing for an AI product covers how to keep your own credits and a third-party allowance cleanly separated in your ledger, which matters once some requests are billed to you and some to the user's plan.
Finally, do not let the plan route make you lazy about efficiency. Every token your app wastes is now visible to the user in ChatGPT's usage view, attributed to your app by name. The same techniques that cut your own bill (prompt caching, routing easy work to a cheaper model, and lowering reasoning effort where it does not help) now cut the share of the user's allowance you consume. We cover them in our guides to prompt caching, model routing and the effort dial. On the plan route, efficiency stops being a cost line and becomes a competitive feature users can see.
8. The Other Ways to Let Users Pay, Compared
Sign in with ChatGPT is the newest and most visible way to move inference cost to the user, but it is not the first and it is not the only one. Each alternative makes a different trade between friction, reach, model freedom and control, and several can coexist in one app. Knowing the field matters for two reasons: some of your users will not have a ChatGPT plan, and a single-lab billing path is a dependency you should be able to replace. This section walks through the six alternatives from the scored table at the top, plus the two labs that have not opened an equivalent door.
Classic bring-your-own-key
The oldest pattern is BYOK: the user creates an account with a model provider, adds billing, generates an API key and pastes it into your app. Cursor's documentation states the economics plainly: the provider bills the user directly for the model cost of those requests - Cursor docs. The same page shows the hidden costs. On Teams and Enterprise plans Cursor adds a $0.25 per million token rate to third-party model requests, the key is sent to Cursor's backend with every request because prompts are assembled there, and Cursor's Zero Data Retention policy does not apply when users bring their own keys.
BYOK's great strength is completeness: the user gets the full API, every hosted tool and every parameter, at list price. Its great weakness is that only technical users will ever do it, and that your app becomes a custodian of long-lived secrets. That custody risk is not theoretical. In February 2026, a misconfigured database at Moltbook exposed about 1.5 million API authentication tokens, and some stored messages contained plaintext OpenAI API keys, according to Wiz's findings - Business Today. Sign in with ChatGPT's short-lived, scoped, revocable tokens are a direct improvement on pasted keys, which is the same argument we made about agents in our guide to giving an AI agent an identity instead of an API key.
OpenRouter OAuth
OpenRouter offers the closest multi-lab equivalent of the ChatGPT button. Users connect their OpenRouter account in one click through PKCE, and your app exchanges the returned code for a user-controlled API key; localhost callbacks work on any port, which suits command-line and local-first tools - OpenRouter docs. The user pays from their own OpenRouter credits, and OpenRouter states that it adds no markup on inference pricing while charging a fee when credits are purchased - OpenRouter FAQ. That fee is 5.5% on the standard tier - OpenRouter pricing.
The trade is clear. OpenRouter gives your users access to many labs' models through one connection, with no approval gate for you as a developer, which makes it the strongest hedge against single-lab dependence. But the user must already have an OpenRouter account funded with credits, and that population is mostly developers. In practice, OpenRouter OAuth and Sign in with ChatGPT complement each other well: one reaches the ChatGPT subscriber who will never open a developer console, the other reaches the power user who wants model choice.
Hugging Face and Puter
Hugging Face lets an OAuth app or a Space request the inference-api scope, which allows it to call Inference Providers on the user's behalf - Hugging Face docs. Usage is billed to the signed-in user at the provider's rates, which Hugging Face says it passes through with no additional fees, and PRO users receive $2.00 of monthly inference credits - Hugging Face pricing. It is the natural choice for apps built on open-weight models, and its reach is concentrated among machine-learning practitioners.
Puter.js takes the user-pays idea furthest for frontend developers. Users sign in to your app with a Puter account, and that account covers the AI, storage and other resources they use, so "their usage never touches your bill"; every user starts with a free monthly allowance - Puter docs. The appeal is that a static site with no backend can ship AI features at zero cost to the developer. The cost is that your users must trust and adopt a brand most of them have never heard of, and your app's capability is bounded by what Puter chooses to offer.
On-device inference
The cheapest inference is the kind nobody pays for. Apple's Foundation Models framework lets apps call the on-device model behind Apple Intelligence, using "AI inference that is free of cost" - Apple Newsroom. At WWDC 2026 Apple went further: developers in the App Store Small Business Program with fewer than 2 million first-time downloads can use the next generation of Apple Foundation Models running on Private Cloud Compute at no cloud API cost, and the same API can call models such as Claude and Gemini from providers that implement Apple's new language model protocol - Apple Newsroom.
On-device inference is the right default for small, frequent, privacy-sensitive tasks such as summarizing a note, classifying an email or extracting fields from a form. It is the wrong tool for the long, tool-heavy reasoning that coding agents and research assistants do. The practical architecture for many apps is a tier: on-device for the cheap and private, the user's plan for the heavy and personal, and your own key for the features no other path supports. If you are weighing self-hosting as a fourth tier, our guide to the best open-weight models to self-host covers the cost side.
Anthropic: the door is closed
Developers who build on Claude should know that Anthropic has taken the opposite position. Its Claude Code legal page states that Anthropic does not permit third-party developers to offer Claude.ai login in their own applications or to route requests through Free, Pro or Max plan credentials on behalf of their users; the permitted case is an end user signing in to the unmodified Claude Code binary with their own subscription - Claude Code legal and compliance. In February 2026, Anthropic clarified that OAuth tokens from consumer Claude plans may not be used in any other product, and OpenCode removed support for Claude Pro and Max keys citing Anthropic legal requests - The Register.
The policy has kept moving since. In April, subscribers were told they could no longer use their Claude subscription limits for third-party harnesses including OpenClaw - TechCrunch. In May, Anthropic announced a separate usage pool for programmatic and third-party use funded by a monthly credit - The Register. Then, on June 15, the day it was due, Anthropic paused that change; its support article says Agent SDK, claude -p and third-party app usage still draw from subscription limits while it works on a plan to better support how users build with Claude subscriptions - Claude Help Center. The lesson for builders is not about Anthropic specifically: subscription allowances are a policy, not a protocol, and policies change quickly.
Google: a first-party door only
Google's subscribers can use their plan in Google's own developer tools. Gemini CLI's "Log in with Google" path allows up to 1,000 requests per user per day on Gemini Code Assist for individuals, 1,500 on Google AI Pro and 2,000 on Google AI Ultra - Gemini CLI docs. There is no published equivalent that lets a third-party app bill Gemini usage to a user's Google subscription. Google has also shown it will act against third-party use of subscription backends: in February 2026 it restricted some Antigravity and Gemini AI Ultra users tied to OpenClaw, citing a "massive increase in malicious usage" of the Antigravity backend - Business Today.
Seen together, the three largest labs have taken three distinct stances. OpenAI has published a formal contract for third-party plan usage. Anthropic forbids third-party login with consumer plans while tolerating some Agent SDK use. Google keeps subscription usage inside its own tools. That divergence is the strongest argument for treating any single plan route as one billing path among several, which the decision framework in Section 12 builds on.
9. Who Has Shipped It, and How They Price It
The launch partners are worth studying not because they are famous but because each made a concrete product decision that newer apps can learn from. The pattern across them is striking: nearly every app that offers plan usage at launch is an agentic tool that runs work for the user, usually around code, and most present the ChatGPT plan as one billing source among several rather than the only one. This section looks at what a few of them actually shipped, using their own announcements and reporting.
Warp added Sign in with ChatGPT to the Warp Terminal and Warp Agent CLI on launch day, letting users "use your subscription's included Work and Codex usage for AI requests in Warp", with usage visible and capped from ChatGPT settings - Warp blog. Kilo Code went further in scope: developers can use up to 100% of the Work and Codex usage included in their plan inside Kilo, and the plan now reaches Kilo's Cloud Agents and mobile app, which previously could not use a ChatGPT plan because requests from Kilo's own infrastructure had no way to reach the subscription - Anaconda blog.
Kilo's implementation is the clearest illustration of how a mature product slots the feature in. The screenshot below, from Kilo's announcement, shows the ChatGPT subscription offered as a provider on the same settings page as classic bring-your-own-key API keys, while the same model menu also lists the 500-plus models in Kilo's own gateway.
The design choice visible here is the one most apps should copy. The ChatGPT plan is not bolted on as a special mode; it is a provider, alongside API keys and the vendor's own credits. That keeps the billing decision with the user, keeps every model reachable through some path, and means a change in OpenAI's policy disables one row of a settings page rather than the product.
Dactyl, the mobile-app builder from Node.js creator Ryan Dahl, showcased the integration in late August, before DevDay. At that time it limited plan sharing to personal ChatGPT Pro subscribers on its Builder plan or higher; Plus and Business users could sign in but kept consuming Dactyl credits - The New Stack. That staged rollout is a sensible template: start with the plan tier whose limits can actually sustain your workload, then widen.
On the open-source side, OpenClaw added a dedicated siwc login method that uses the Codex allowance with its own usage tracking and token limits per OpenClaw instance, while keeping the older Codex login and API-key options alongside it, and OpenCode exposes it as a "ChatGPT Pro/Plus (browser)" choice with a headless variant for servers - Flavio Copes. The headless variant matters: it shows how open-source tools are solving the self-hosted VM case that OpenAI's docs cover with a credential-transfer procedure.
Two takeaways generalize. First, every serious implementation keeps at least one other way to pay, because users without an eligible plan, users who hit their cap, and features the plan route does not support all need somewhere to go. Second, the apps that benefit most are those whose cost is dominated by model calls the user directly asks for. If your costs are dominated by things the plan route cannot pay for (hosted retrieval, image generation, compute you run yourself), the feature moves a smaller part of your bill, which is exactly the split Dactyl's credits make explicit. Our comparison of Claude Code, Codex and Devin covers how the agent tools in this list differ in what they actually do.
10. Where It Fails: Risks, Abuse and Platform Dependence
Every billing path has failure modes, and the plan route's are unusual because they sit partly outside your app. Some are technical (a stream that dies halfway), some are commercial (a policy that changes), and some are human (a user who thinks signing out stopped the spending). Working through them before launch is cheaper than discovering them in support tickets, and several of them have straightforward mitigations if you design for them from the start.
Platform dependence and single-lab lock-in
The largest risk is structural. The plan route only pays for OpenAI models, under rules OpenAI sets and can change, and commercial access is currently at OpenAI's discretion. If a meaningful share of your inference moves onto users' ChatGPT plans, a future change to eligibility, caps or the trial program changes your gross margin overnight, and you will have trained your users to expect not to pay you for AI. Section 8 showed how quickly another lab's subscription policy moved in a single half-year, through a block, a clarification, a new credit and a pause.
The mitigation is architectural, not contractual. Keep inference behind an internal interface that can route a request to the user's plan, your own key, or another provider, and keep at least one non-OpenAI path working in production even if it serves few users. That is the same independence argument we made in our guide to owning your AI stack: control comes from being able to switch, not from any single provider's goodwill. A billing path you cannot turn off without breaking your product is not a feature; it is a dependency.
Mid-stream failures and silent disconnections
The plan route introduces failures your API-key path never had. A usage-limit error can arrive as response.failed after the answer has started streaming, so your UI can show half an answer that will never finish. OpenAI does not notify your app when a user disconnects it in ChatGPT settings, so you discover a revoked grant only when the next request or refresh fails - OpenAI errors guide. And on Plus, another app can drain the shared five-hour window, so your request can fail even though your app used little.
Design for those states explicitly. Store partial outputs so a user can resume after switching billing paths; show a specific "your ChatGPT plan limit was reached" state with the Manage usage action rather than a generic error; and for long agent tasks, checkpoint work so a limit mid-task costs a pause, not a restart. The apps that handle this well will feel more reliable than the plan itself, which is a real advantage when users compare tools inside the same allowance.
Users who misunderstand what they granted
The user-side controls are good but unfamiliar, and two misunderstandings are already documented. Signing out of an app does not disconnect it from ChatGPT or stop its authorized use of the plan; disconnection happens in ChatGPT's settings - OpenAI Help Center. And a user who allows apps to use credits while automatic credit purchases are on can be charged without a separate notice when included usage runs out. WorkOS sketched the worst case: an application's retry loop running against a grant one engineer made on a personal plan, triggering a card charge "that nobody budgeted for and nobody was told about" - WorkOS.
WorkOS also raised a governance problem that B2B founders should take seriously. An engineer whose employer runs a locked-down Enterprise ChatGPT deployment can sign in to a coding agent with their personal Plus account, and the employer's admin console has no view of that grant even though the work happens on company code. If you sell to companies, decide whether you want work accounts to use personal plans at all, and say so in your docs. The safest default for enterprise customers is identity sign-in with their managed workspace plus billing on your own contract.
Abuse, fraud and the free-trial problem
On the other side of the ledger, the plan route solves a problem that has been expensive for AI startups. Stripe's data shows AI startups faced an attempted transaction fraud rate 4.3 times that of startups overall in Q3 2025, easing to 2.6 times by Q1 2026, and a 40% increase in attempted multi-account abuse across AI subscription companies from January to June 2026, where fraudsters create many accounts "to repeatedly claim free tokens, trials, and other new account benefits" - Stripe. An earlier Stripe analysis found 7.4% of sign-ups at AI companies implicated in suspected multi-account abuse - Stripe.
The chart explains a quiet benefit of the plan route. A free trial funded by your API key is a pile of free tokens that attracts exactly the multi-account abuse Stripe describes. A trial funded by the user's own ChatGPT plan offers nothing to steal, because every request draws down an account that someone already paid for and that OpenAI already verified. That does not eliminate abuse of your other resources, but it removes the most attractive target. For the rest, our pre-launch security checklist covers rate limits, sign-up controls and the other defenses an AI app needs.
Compliance questions to ask before launch
Finally, the plan route changes whose account a request runs under, and that has compliance implications worth checking with counsel rather than assuming. Your organization's API data controls are configured on your API organization; a request made with the user's own ChatGPT token is made under the user's account. If you operate in the EU or sell into regulated industries, ask how your privacy policy and records of processing describe requests that run on a user's own OpenAI account. Region matters too: OpenAI lists serving-region policy among the reasons the plan route can refuse admission with a 403. Our guide to making an AI app EU-compliant by December 2026 covers the obligations that apply regardless of who pays.
11. Agents Change the Math
Everything above gets sharper as apps become agents. A chat assistant spends a few thousand tokens per exchange; an agent that plans, reads files, runs tools and checks its own work can spend millions per task. That is why the per-user cost chart in Section 7 jumps by two orders of magnitude between the light assistant and the heavy agent, and why the launch partners are overwhelmingly agent tools. For agent builders, who pays for inference is no longer a margin detail; it decides whether the business model works at all.
OpenAI's contract is visibly designed with agents in mind. The open-source flow registers a "user-defined agent", distinguishes clients from hosts so one person's laptop and self-hosted VM can share limits while being tracked separately, and documents how to move credentials to a remote VM; OpenAI notes that host-specific usage attribution and revocation for transferred sessions are not yet available - OpenAI self-hosted VM guide. The supported tool model (function and custom tools run by the client, no hosted execution) is exactly how local agents already work, and Codex app-server can run shell and MCP tools locally while billing the user's plan.
The pattern of a user's own plan paying for an agent's work is also older than this launch. Coding CLIs have run on their makers' consumer subscriptions since 2025, when Codex added Sign in with ChatGPT to its IDE extension and CLI - ChatGPT changelog. Founden's desktop app ships its builds the same way: it runs Claude Code or Codex on the founder's own Mac against the founder's own Claude or ChatGPT subscription, and strips vendor API keys from the coder's environment so a build cannot silently fall back to API billing. That design keeps an always-working AI builder from turning into an open-ended token bill, and it is the same economic logic OpenAI has now opened to every open-source tool: the user's plan pays for the work they ask for.
There is a tension here that agent builders should face honestly. Plans are priced for human-paced use. An agent that runs while the user sleeps consumes allowance at machine pace, and the five-hour and weekly limits were not designed for fleets of agents. OpenAI's own always-on dots show how the company itself treats this: conversations with a dot do not count toward ChatGPT limits, but tasks a dot starts in Codex or ChatGPT Work do - The Next Web. Expect plan limits to keep evolving toward agent workloads, probably through more tiers like Pro 500 rather than more generous Plus allowances.
What happens next
Three developments look likely, each grounded in what has already happened rather than in hope. First, the commercial trial will widen: OpenAI told The New Stack it is "starting with these 16 launch partners today, and we'll expand quickly" - The New Stack. Second, competitors will respond with their own allowance portability, because the shape is cheap to copy and Anthropic has already said it is working on a plan to better support building with Claude subscriptions. Third, the identity layer will become a battleground: WorkOS predicted that once one vendor shows an allowance can travel across sixteen products, "the pattern gets copied" and identity teams end up managing a scope whose cost is measured in dollars.
If those play out, the end state resembles the mobile data analogy from Section 2 more closely: users will hold AI allowances from one or more labs, apps will request permission to draw on them, and the app's price will reflect what it adds on top. The builders who do well in that world are those whose product is worth paying for even when the AI is free to them. If your product sells inside ChatGPT itself, our guide to selling your product inside ChatGPT covers the commerce side of the same shift.
12. A Decision Framework for Your App
With the mechanics, the money and the risks laid out, the decision for a specific app comes down to a handful of questions asked in the right order. The order matters because the first questions are hard gates (can you use the feature at all?) and the later ones are economic judgments (should you?). The diagram below captures the sequence, and the paragraphs after it explain each branch.
The first gate is eligibility. If your app is open source or runs on the user's own machine, the self-serve flow is open to you today, and the cost of adding it is a few days of OAuth work. If you run a closed, hosted product, apply through OpenAI's interest form and build against the published contract in the meantime, so you can ship the day you are approved. Either way, ship identity sign-in where you can: a recognizable login button lowers sign-up friction even before plan usage is enabled.
The second gate is capability. If your product's core value depends on OpenAI-hosted tools, sampling controls or server-side conversation state, the plan route cannot carry those calls, so keep them on your own key and route only the compatible requests to the user's plan. This split architecture is more work, but it is also the architecture that survives policy changes, because each request already knows which billing paths can serve it. Remember too that the plan route serves only OpenAI models; if part of your product runs better on another lab's model, that part stays on your own bill or a multi-lab path like OpenRouter.
The economic question comes last. Measure your inference spend per user, find your heaviest users, and estimate how many of them already pay OpenAI. If your bill is concentrated among people who are likely Plus or Pro subscribers (developers, analysts, writers, researchers), a bring-your-plan tier can move a large share of your cost off your books. If your users are mostly casual consumers or mostly enterprise employees on managed workspaces, the plan route reaches fewer of them, and identity sign-in plus your own pricing is the better default for now. Whichever branch you land on, budget for the alternatives your users will need: API-key billing, your own credits, and possibly on-device inference for the light, private work.
For founders building their first AI product, the broader cost picture (hosting, databases, models and the rest of the stack) is in our guide to what it costs to build an app with AI, and the login decision itself, including where a ChatGPT button sits next to email and social sign-in, is covered in our comparison of authentication options for your app.
Conclusion
Sign in with ChatGPT's plan usage is the first formal, documented way for a major AI lab's subscribers to carry their allowance into other companies' apps. For the right app, it is a genuine shift: an eligible user can start using your AI with one click, no API key and no card, while the inference cost settles between them and OpenAI. For open-source and local tools, especially agents, it is available today and worth implementing now. For hosted commercial products, it is a trial worth applying to and a contract worth building against.
It is also narrower than the keynote made it sound. The plan route serves only OpenAI models, strips out hosted tools and sampling controls, puts the meter in the user's hands, shares a five-hour window across apps on Plus, and depends on a policy that a single company can change. The labs have already diverged: OpenAI opened the door, Anthropic closed its equivalent for third-party login, and Google kept its subscriptions inside its own tools. Those facts point to one durable strategy rather than a bet on any single button.
| If your app is... | Do this now |
|---|---|
| Open-source or local agent | Implement plan usage now; keep API keys as a fallback |
| Hosted, agent-heavy product | Apply, build the split architecture, price a bring-your-plan tier |
| Hosted, hosted-tool-heavy product | Identity sign-in now; plan route only for compatible calls |
| Sold to enterprises | Managed identity and your own billing, not personal plans |
| Light, private features | Consider on-device inference before any paid path |
The common thread in every row is independence. Treat the user's ChatGPT plan as one billing provider among several, route every request through an interface you control, and keep at least one path that does not depend on OpenAI's approval or OpenAI's models. Then make sure your product is worth paying for even when the AI is not on your bill, because that is the world this feature points toward: users holding their own AI allowances and paying apps for what they add on top. Builders who do that get the upside of OpenAI's new distribution without handing it the keys to their margin.
This guide reflects Sign in with ChatGPT, ChatGPT plan pricing and the related AI pricing as of October 7, 2026. The feature is labelled a preview, plan eligibility differs between OpenAI's own pages, and limits, prices and partner lists change frequently, so verify current details in OpenAI's documentation before building on them.