Which AI app builders and coding tools train on your code in 2026, what changed this autumn, and exactly how to switch it off
Since September 9, 2026, Lovable may use everything you build on its Free and Pro plans to train its AI models. That covers your prompts, the files you attach, the code it writes, your project files, your configurations and even your hosted app, unless you find one toggle and turn it off - Lovable Docs. Opting out is free, but it only works going forward, and it only covers your own account: every collaborator in a shared workspace has to flip their own switch.
Lovable is not an outlier. It is the latest step in a wave. GitHub began training on interaction data from Copilot Free, Pro and Pro+ users on April 24, 2026 - GitHub Changelog. Vercel opted Hobby accounts into model training for v0 in March - Vercel Changelog. And Bolt published new terms on September 14 that, from October 7, 2026, let it train on individual accounts by default and license datasets derived from your projects to third parties, "including for compensation" - StackBlitz Terms of Service. If you build software with AI, the question "does this tool train on my code?" now has a different answer than it did a year ago, and in most cases the answer depends on which plan you pay for and whether you have ever opened your settings.
Here is the problem for founders: the toggles are scattered, they are named differently in every product, they rarely apply retroactively, and switching one off does nothing to the others. A founder who prototypes in Lovable, iterates in Cursor and fixes bugs with Claude Code has at least three separate decisions to make, and a team has those decisions multiplied by every person on it. Meanwhile the thing at stake is not abstract. Your prompts describe your business model, your code encodes your product, and your chat history often contains the API keys and customer details you pasted in at 2am.
This guide answers the search question directly (does Lovable train on your data, and how do you stop it), then goes much further. It explains why AI builders want your code in the first place, what "training" actually means in a privacy policy, and where every major builder and coding assistant stands today, with the exact setting to change and the vendor's own wording as the source. It covers the levers beyond a settings toggle (law, contracts, acquisitions, crawlers), gives you a 20-minute audit you can run today, and is honest about what opt-outs cannot do.
Contents
- The Short Answer: Does Lovable Train on Your Data?
- Why AI Builders Want Your Code in the First Place
- What "Training on Your Data" Actually Means
- Lovable in Depth: What Changed on September 9
- Bolt, v0, Replit, Base44 and Figma: The Other App Builders
- Coding Assistants: Copilot, Cursor, Claude Code, Codex and More
- The Model Layer: APIs, Subscriptions and Retention
- The Price of Default Privacy
- Beyond the Toggle: Law, Contracts and Acquisitions
- Protecting What You Publish: AI Crawlers and Your Live App
- The 20-Minute Opt-Out Audit
- What Opt-Outs Do Not Do
- Where This Is Heading
- Conclusion: A Decision Framework
How the major AI builders score on protecting your code
The table below scores thirteen AI app builders and coding assistants on how well they keep your code out of model training, using only each vendor's current published policy. It is sorted by final score. Founden, which publishes this guide, is scored on the same criteria as everyone else and sits where its score puts it, gaps included.
| # | Tool | What It Does | Default on Entry Plans (35%) | Opt-Out Quality (25%) | After You Opt Out (20%) | Contract Strength (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | JetBrains AI | AI assistant and agent inside JetBrains IDEs | 7 - paid licenses share only if you opt in; free non-commercial licenses share by default | 8 - one "detailed data collection" setting | 6 - shared data kept up to one year | 6 - stated in a public FAQ, not a DPA | 6.9 |
| 2 | Founden (our product) | Builds autonomous software that runs as a company | 6 - Founden trains no models; desktop builds inherit your own Claude or ChatGPT plan's toggle | 7 - nothing to switch off at Founden, but your AI plan's toggle is yours to set | 6 - prompts, build steps and results are stored by Founden | 7 - terms license your content "solely as necessary to provide the Services"; no public DPA | 6.5 |
| 3 | Cursor | AI code editor, now owned by SpaceX | 5 - Teams get Privacy Mode by default; individuals must check it | 8 - one Privacy Mode toggle on any plan, zero retention at providers | 7 - no training with Privacy Mode on; abuse and own-key exceptions | 6 - data-use page updated Sep 3, 2026 after the ownership change | 6.4 |
| 4 | Claude Code | Anthropic's terminal coding agent | 4 - Free, Pro and Max sessions train when "Help improve Claude" is on | 8 - one toggle covers chat and Claude Code | 6 - not retroactive; feedback transcripts kept 5 years | 8 - commercial terms bar training; Team is off by default | 6.2 |
| 5 | OpenAI Codex | OpenAI's coding agent, in ChatGPT and CLI | 4 - personal ChatGPT plans train by default, Codex included | 7 - one toggle, but Codex has a separate environments setting | 5 - feedback overrides the opt-out | 8 - Business, Enterprise and API off by default | 5.8 |
| 6 | GitHub Copilot | AI pair programmer inside GitHub and IDEs | 3 - Free, Pro, Pro+ train since Apr 24, 2026 | 8 - one Privacy setting; earlier opt-outs kept | 5 - not retroactive; shared with affiliates incl. Microsoft | 8 - Business and Enterprise excluded by contract | 5.7 |
| 7 | v0 (Vercel) | Prompt-to-app builder on Vercel | 5 - Hobby and Trial Pro opted in; paid Pro opted out | 7 - Team Settings, Data Preferences, per-project override | 4 - code may go to model providers; "can't be unshared" | 7 - in the terms; Enterprise never trained | 5.7 |
| 8 | Amazon Kiro | AWS agentic IDE | 4 - Free Tier and individual subscribers on by default | 7 - one checkbox under Telemetry and Content | 5 - Free Tier inputs kept up to 60 days | 7 - enterprise identity users excluded | 5.6 |
| 9 | Lovable | Prompt-to-app builder, full stack | 3 - Free and Pro train by default since Sep 9, 2026 | 6 - free on any plan, but per person only | 3 - not retroactive; de-identified data for "any lawful purpose" | 6 - Business and Enterprise covered by a DPA | 4.4 |
| 10 | Google Antigravity | Google's agentic IDE, successor for Gemini CLI users | 3 - personal accounts: interactions improve ML, human review | 4 - "navigate to settings"; users report only a telemetry switch | 4 - deletion by email request | 7 - Workspace, Enterprise and paid API excluded | 4.3 |
| 11 | Bolt | Prompt-to-app builder by StackBlitz | 2 - individual accounts outside EU, UK, Switzerland on from Oct 7, 2026 | 5 - free on any plan, except Bolt Forge content | 2 - datasets may be licensed to third parties | 5 - Teams and Enterprise excluded | 3.4 |
| 12 | Replit | Browser IDE and app-building agent | 3 - policy claims a legitimate interest in improving code generation | 2 - no documented opt-out setting found | 3 - private apps accessible to improve the service | 6 - Enterprise agreement bars training | 3.4 |
| 13 | Base44 | Prompt-to-app builder owned by Wix | 1 - every plan except Enterprise "can be used to train" | 1 - no opt-out found below Enterprise | 3 - retention after training not documented | 5 - Enterprise excluded automatically | 2.2 |
How to read the criteria. Default on entry plans (35%) carries the most weight because, as Section 2 explains, defaults decide what actually happens to most users' data. It scores the plans a founder realistically starts on: free and the first paid tier. Opt-out quality (25%) asks whether anyone on any plan can opt out, for free, in one obvious place, covering all their content. After you opt out (20%) scores what still happens: retention windows, retroactivity, human review, dataset licensing and other uses that survive the toggle. Contract strength (20%) scores whether the commitment is binding (terms of service, a data processing agreement) rather than a help-center sentence, and how affordable the protected tier is.
Two notes on the scores. First, a high score does not mean "never trains"; it means the vendor's defaults and commitments leave you less exposed with less effort. Second, the scores apply to each vendor's published policy on October 3, 2026, and several of these policies changed within the last six months. The sections below give the exact wording behind every cell.
1. The Short Answer: Does Lovable Train on Your Data?
Yes, by default, if you are on the Free or Pro plan. Lovable's documentation states that "as of September 9, 2026, customer data from Free and Pro plans may be used to train, develop, and improve Lovable's AI models and AI-powered features." Customer data is defined broadly: "prompts (including images and files you attach), code, project files, hosted applications, configurations, and generated outputs" - Lovable Docs. Business and Enterprise workspaces are "excluded from model training by default" and governed by their data processing agreement, so no action is needed there. The data your app's own end users type into it, such as sign-ups and form submissions in your app's database, is not used for training.
Two details change how you should act on that answer. The opt-out is per account, not per workspace: "It does not change the setting of other members in a shared project or workspace. Each person controls their own." And it is not retroactive: Lovable's privacy policy says opting out "does not retract content from training datasets assembled, or models trained, before you opted out" - Lovable Privacy Policy. If you opted out before September 9, the original change notice confirms your data was never used.
To stop Lovable training on your data:
- Open Account settings while signed in, then Preferences, then AI model training (the direct link is lovable.dev/settings/account#privacy).
- Turn off the toggle labelled "Use my Lovable content for model training."
- Ask every collaborator on a Free or Pro workspace to do the same on their own account.
If you cannot sign in, Lovable accepts the same choice by email at privacy@lovable.dev. The steps take about a minute, and they are free on every plan: Lovable says you can opt out "at any time in your account settings, on any plan, at no cost," and that doing so "does not affect your use of AI features" - Lovable Privacy Policy. There is no capability penalty and no reason to delay. The only cost is remembering to repeat the check for every person who contributes to a project, because Lovable gives Free and Pro workspaces no admin switch that covers everyone at once.
The label is worth a second look. Lovable's August change notice described the control as "Account Settings → Privacy" with a "Data collection opt out" switch - Lovable Summary of Changes, while the current documentation calls it "AI model training" under Preferences. If your screen shows either wording, you are in the right place; if it shows neither, the email route works on every plan. And if your project is commercially sensitive enough that a missed toggle on one collaborator's account would worry you, the structural answer is a plan where training is off by default, which on Lovable is Business at $50 per month - Lovable Pricing. Section 8 compares what that guarantee costs across every major tool.
What about the projects you built before you found the setting? If you opted out before September 9, nothing changes: Lovable's notice says "if you opt out before the effective date, your data is not used for training" - Lovable Summary of Changes. If you opt out now, everything you created since September 9 may already be in a training dataset, and deleting a project does not pull it back out, because the privacy policy says deletion "does not retract content from training datasets assembled, or models trained, before your request took effect." The realistic damage control is not deletion but hygiene: rotate any API key or password that ever appeared in a prompt, as Step 4 of the audit in Section 11 explains.
Should you stop using Lovable because of this? Not necessarily. The change is disclosed, the opt-out is free and immediate, and the policy is more detailed about its own limits than many competitors' are. Several of the alternatives in Section 5 are less protective, not more. The right response for most founders is to switch the setting off today, decide whether the project justifies a Business plan, and treat the episode as a reminder that every AI builder's default now deserves the same one-minute check.
2. Why AI Builders Want Your Code in the First Place
The usual explanation for why AI companies want your data is that "data is the new oil," which explains nothing. The structural reason is narrower and more useful. Public code has already been harvested. Every serious coding model has seen the open repositories, the documentation and the Stack Overflow archive. What none of them can find on the open web is the process of building: a plain-language request, the plan the agent made, the code it wrote, the error that came back, the fix, and above all the human verdict at the end. Did the founder keep the change, undo it, ask again, or ship it? That sequence (the industry calls it a trajectory) is the training signal that teaches an agent to build working software rather than plausible-looking software, and it only exists inside the products where the building happens.
Cursor described this loop in public. Its Tab model "runs on every user action, handling over 400 million requests per day," and the company says it has "a lot of data about which suggestions users accept and reject," which it used to retrain the model with online reinforcement learning, rolling out a fresh checkpoint every 1.5 to 2 hours - Cursor. The result was a model that made 21% fewer suggestions with a 28% higher accept rate. A month later Cursor shipped Composer, its own coding model, specialised "through reinforcement learning (RL) in a diverse range of development environments" - Cursor. Cursor does not say Composer was trained on customer sessions, and you should not assume it was. The point is the direction of travel: AI builders are no longer thin wrappers around a lab's model. They increasingly train models of their own, and the cheapest, richest training material they could possibly have is what their own users do every day.
The free tier is paid in data
Every message you send to an AI app builder costs the builder real money, because it pays a model provider for each token the agent reads and writes. We broke down those token meters in our guide to what it costs to build an app with AI. A free plan that hands out daily credits is therefore a subsidy, and subsidies get recovered somewhere. Historically that was conversion to a paid plan. Training rights are a second recovery channel, and an attractive one, because the value of a session's data does not depend on whether the user ever pays.
This is why the line between "trains by default" and "does not train by default" so often falls along plan boundaries rather than along any principle about the data itself. The same prompt, written by the same person, gets treated differently depending on whether they pay $0, $25 or a business price. When you read a builder's training policy, read it as a pricing page. It tells you what your data is worth to the company and how much you would have to pay to keep it.
Why the default matters more than the toggle
Almost every builder that trains on user content offers a way out, and that fact gets presented as if it settles the matter. It does not, because defaults decide outcomes. The classic evidence comes from organ donation: in Eric Johnson and Daniel Goldstein's 2003 study, European countries where citizens had to opt in averaged around 10% consent, while countries where they had to opt out sat close to 100% - Kellogg School of Management. Same kind of people, same stakes, different checkbox.
Apply that to an AI builder. A company that flips its default from "off" to "on" does not gain a trickle of volunteers. It gains, in practice, most of its free and entry-tier users overnight, because most people never open the setting. That is why the date a default changes is the event that matters, and why this guide spends as much time on dates and defaults as on the toggles themselves.
Training is a one-way door
The final structural fact is an asymmetry. You can switch a setting off in a second. A model cannot reliably forget what it learned from your project, and the vendors' own policies are written accordingly: opting out stops future use, while content already folded into a trained model stays there. Anthropic's privacy center says it plainly: after you change the setting, "your data will still be included in model training that has already started and in models that have already been trained" - Anthropic Privacy Center. So the cost of a late opt-out is permanent, and the cost of an early opt-out is close to zero. That asymmetry is the whole argument for doing the audit in Section 11 today rather than "when the product matters."
It also explains why memorization is not a theoretical concern. In a peer-reviewed study published at FSE 2024, researchers prompted GitHub Copilot with real code whose secrets had been removed. Out of 8,127 suggestions, Copilot produced 2,702 plausible credentials, and 200 of them matched real secrets from GitHub code exactly - arXiv. That study concerned public code, not private projects, but the mechanism is the same one that makes a training default worth a founder's attention: what goes into a model can come back out, verbatim, for somebody else.
3. What "Training on Your Data" Actually Means
"We use your data to improve our services" and "we train AI models on your content" sound similar and mean different things. Before you can opt out of anything, you need to know which of several distinct activities a policy actually covers, because a single toggle rarely switches all of them off. The vocabulary below is how the policies themselves are written, and it is worth learning once.
None of this vocabulary is exotic, and none of it requires legal training to use. It simply gives you the right questions to ask of any policy, including ones written after this guide.
The first distinction is between content and telemetry. Content is what you create: prompts, chat history, generated code, uploaded files, images, the database schema the agent designed. Telemetry is how you use the product: clicks, feature usage, error rates, latency, which suggestion you accepted. Many policies exclude content from training while keeping telemetry, and a few treat "which suggestion you accepted" as telemetry even though, as Cursor's Tab work shows, that signal is exactly what reinforcement learning feeds on.
The five things a policy can mean
Read any AI builder's privacy policy or terms and you will find some combination of five activities. They differ in who sees your content, for how long, and whether it ends up inside a model's weights, and most policies switch them on and off independently. The order below runs roughly from least to most consequential for a founder who cares about their code, which is also roughly the order in which vendors make them hard to avoid.
The reason this matters in practice is that a single sentence in a policy can quietly cover several of these at once. "We use your content to improve our services" can mean retention plus human review plus model training, while "we do not train on your content" can still leave retention and human review in place. Knowing the five categories lets you read past the headline and ask which ones a specific toggle actually controls.
- Processing to serve you: sending your prompt to a model so it can answer. Unavoidable, and covered by every terms of service.
- Retention: keeping your content for a period (30 days, 5 years, indefinitely) after the session, often for abuse monitoring or support.
- Human review: staff or contractors reading sampled conversations, usually for safety or quality labelling.
- Service improvement: using aggregated or de-identified data to tune product behaviour, evaluate releases or fix bugs.
- Model training: using your content to update the weights of a model, whether the builder's own or a partner's.
The practical lesson from this list is that "opt out of training" usually switches off only the last item, and sometimes the fourth. Retention and human review often continue under separate rules, and processing never stops while you use the product. So when a vendor tells you that you "can opt out," the follow-up questions are: out of which of these, starting when, and does it touch what you already submitted. The platform sections below answer those questions tool by tool, citing the vendor's own wording rather than paraphrasing it.
A worked example makes the difference concrete. Suppose you paste a failing database query into a builder's chat, along with the error message and a sample customer record, and the agent fixes it. With training switched off, that exchange is still processed (the model has to read it), still retained for the vendor's standard window, possibly reviewed if it trips an abuse filter, and possibly counted in aggregate analytics. What the toggle prevents is the last step: the exchange being folded into the dataset that updates a model. That is the step that cannot be undone, which is why it gets the toggle, but it is not the only thing that happens to your text.
The same example shows why the category "telemetry" deserves suspicion. Whether you accepted the agent's fix, how long you looked at it, and whether you asked for another attempt are behavioural signals rather than content. Several policies treat them as usage data rather than customer content, yet they are exactly the reward signal that reinforcement learning needs. When a policy promises not to train on your "content" but says nothing about "usage data," assume the behavioural signal is still flowing.
Two layers: the builder and the model behind it
The second distinction is structural, and it is the one most founders miss. When you type into an AI app builder, your content can be exposed to training at two separate layers. The first is the builder itself: Lovable, Bolt, Replit or Cursor deciding what to do with your project. The second is the model provider the builder calls: Anthropic, OpenAI, Google and others. The two layers have separate contracts, separate defaults and separate opt-outs, and switching off one does nothing to the other.
In practice, builders usually reach models through the providers' commercial APIs, which (as Section 7 shows) do not train on inputs by default. That makes the builder layer the one that changes under your feet, and the Lovable change is a builder-layer change. The exception is the newer pattern where an AI coder runs on your own consumer subscription, such as Claude Code on a Claude Pro plan or Codex on ChatGPT Plus. There, the model provider's consumer data controls apply to you directly, and the builder cannot set them for you.
The two layers also fail in different ways. A builder-layer change arrives as an email from a company you chose and pay, with a few weeks' notice, and affects one product. A model-layer change affects every tool that runs on that provider's consumer plan at once, because the setting lives on your account with the provider rather than inside any single tool. Keeping both layers in view is what lets you answer the question this guide is named after for any tool, including ones that did not exist when it was written.
The diagram explains why one setting is never enough. A founder who builds in a web builder and also runs an AI coder on a personal subscription has at least two toggles to find, and a team has one per workspace and one per person. Section 11 turns this into a checklist you can run in about twenty minutes.
4. Lovable in Depth: What Changed on September 9
Lovable announced the change on August 5, 2026, with a summary notice that gave users five weeks before it took effect on September 9 - Lovable Summary of Changes. It then revised the full privacy policy again, with an effective date of September 15, 2026 - Lovable Privacy Policy. For anyone searching "does Lovable train on my data," the short answer in Section 1 is enough to act on. This section is for founders who want to understand exactly what they agreed to, because the details matter more than the headline.
The single most important sentence is the scope of training. Lovable says it uses your customer content and usage data "to train, develop, fine-tune, and improve our AI models and AI-powered features, including models we operate within the Services and models we may make available to customers through Lovable products such as the AI Gateway." That last clause is the strategic tell. Lovable is not only tuning the assistant you use. It is reserving the right to train models it can offer to other customers, which means the patterns in your project can, in principle, inform tools your competitors use too.
What the policy covers, and on what legal basis
Lovable's notice states the legal basis directly: "We will rely on legitimate interests to do this." That is a meaningful choice. Under European data protection law, consent has to be asked for and can be refused up front, while legitimate interest lets a company start processing and gives you a right to object instead (Section 9 explains what that right does and does not get you). In plain terms, Lovable chose the legal route that makes training the default rather than the exception, which is consistent with how the product now behaves.
The policy also documents two practices that survive into the training pipeline. First, human review: "Trained members of our team may review this content to check model quality and diagnose failures." Second, de-identified data: Lovable "may use and disclose De-identified Data for any lawful purpose," and its documentation separately says it may create de-identified data and use it "for purposes such as improving the Services." De-identification is a real safeguard for personal data, but it does little for a founder whose concern is the product idea, the architecture or the prompt strategy, because none of those identify a person in the first place.
There are things Lovable explicitly does not train on, and they are worth knowing so you do not over-correct:
- End-user data your app's visitors submit, which stays in your project's own database.
- Account and billing details, used only to run your account.
- Business and Enterprise workspace content, excluded under a data processing agreement.
- Messaging integration data from a collaboration platform you connect to Lovable.
The practical reading of that list is that Lovable has drawn its line around what you create rather than around what your customers create. That is the right line from a privacy-law perspective, because your customers never agreed to Lovable's terms, but it is the opposite of where a founder's commercial risk sits. Your customers' personal data was already protected by law; your product's design, prompts and code are protected only by the setting you choose. That asymmetry is the single best reason to check the toggle even if you have nothing personal in your project at all.
The advertising change that arrived the same day
The September 9 update carried a second change that most coverage missed. As of that date, Lovable "may share limited pseudonymized identifiers with advertising platforms." In the EEA, the UK, Switzerland and Brazil this requires consent; in the United States you can opt out through the "Do Not Sell or Share My Personal Information" link, your privacy settings or a Global Privacy Control signal - Lovable Docs. Lovable is explicit that it does not share the contents of your projects with ad platforms.
This is a separate control from model training, and turning one off does not turn off the other. If you are a US founder who wants neither, there are two switches to find. The cleanest method is to enable Global Privacy Control in your browser, which Lovable names as a valid opt-out signal, and then still flip the training toggle, because GPC covers the "sale or sharing" of personal data and says nothing about model training.
The team problem: one toggle per person
The most underappreciated detail in Lovable's design is that the opt-out lives on each account, not on the workspace. On Free and Pro there is no admin setting that applies to everyone. If a founder opts out but a contractor, a co-founder or a friend helping with the design does not, the content that person contributes to the shared project remains eligible for training. On a project with five collaborators, that is five settings to verify, and nothing in the product tells you who has and has not done it.
This matters because AI app builders are increasingly used by small teams rather than solo makers, as we found when ranking the field in our top 20 AI app builders guide. The structural fix is a plan where training is off at the workspace level. On Lovable that means Business, where workspace data is "excluded from model training by default" and handled under the data processing agreement. The cost difference between Pro at $25 per month and Business at $50 per month - Lovable Pricing is, in effect, the price Lovable puts on not having to police five toggles. Whether that is worth paying depends on what the project is, which is exactly the question Section 14's decision framework answers.
Lovable's acquisition of Sutro, and why ownership matters
Nine days after the training change, Lovable announced it had acquired Sutro, a team that "for five years... built a language, compiler, and backend platform that made application behavior and rules explicit" - Lovable Blog. The announcement says nothing about data or training, and there is no reason to think the deal changes Lovable's privacy policy. It is still worth noting for a structural reason: builders are consolidating fast, and every acquisition puts your data under a new owner's incentives.
The image marks one of several builder-stack deals in a few weeks, and Section 9 looks at the pattern. The lesson for Lovable users is narrow but durable: the policy you read today is a statement of current intent, not a permanent contract term, and the August notice is proof that a builder can change the default with five weeks' warning. A setting you chose once should be rechecked every time a vendor emails you about "updates to our privacy policy," because that email is the only warning you will get.
5. Bolt, v0, Replit, Base44 and Figma: The Other App Builders
Lovable's change drew attention because it is one of the most widely used prompt-to-app builders, but its competitors have moved in the same direction this year, some of them further. The pattern across the category is consistent: individual and free accounts are opted in, team and enterprise accounts are opted out, and European accounts sometimes get a different default because European law makes default-on training harder to defend. What differs, and differs a lot, is what happens to your data once it is collected.
If you have been comparing these tools on price and output quality, as we did in our Lovable vs v0 vs Bolt cost breakdown, add data terms as a third axis. A builder that is a few dollars cheaper per app but licenses datasets from your project to third parties is not cheaper for every kind of project.
Bolt: from October 7, your projects can become licensed datasets
Bolt, the builder made by StackBlitz, published new terms on September 14, 2026, updated them on September 22, and set them to take effect for existing accounts on October 7, 2026 - StackBlitz Terms of Service. Its release notes summarise the default: "It's on by default, except for accounts based in the EU, UK, or Switzerland, where it's off by default" - Bolt Release Notes. The data covered is broad and specifically agentic: AI inputs and outputs, "error messages, correction or fix traces, tool invocations, and edit histories," and "project files, code, and configuration."
The clause that sets Bolt apart is dataset licensing. Unless you opt out, you grant StackBlitz "a worldwide, non-exclusive, royalty-free, sublicensable, and transferable license to prepare, market, license, and distribute datasets derived from" your eligible content "to third parties, including for compensation (which some laws may characterize as a 'sale')." Read that with the first-principles lens from Section 2: the trajectories that teach an agent to build working software are valuable enough that a builder can sell them. And the opt-out "does not, by itself, unwind processing completed before the request, models already trained, or datasets already delivered to licensees."
Three timing and scope details decide whether this affects you:
- Content before October 7 is not eligible, unless it is Bolt Forge content.
- Bolt Forge content from September 14 is eligible, and the opt-out does not apply while Forge is on.
- EU, UK and Swiss accounts are excluded, decided from billing address and network location.
- Teams and Enterprise workspaces are excluded entirely.
The Forge exception deserves its own warning. Forge is a research-preview mode included with individual Pro plans until October 14, and the terms say the general opt-out "does not apply to Forge Content while Forge is enabled." To stop that, you have to stop using Forge, and even then "stopping use of Forge does not by itself withdraw your consent for Forge Content already collected." If you build on Bolt today, the safe sequence is: turn off Forge, then open your avatar, Settings, General, and switch off Model training - Bolt Help Center, or email privacy@stackblitz.com. If you want the default handled for the whole team, Bolt Teams costs $30 per member per month - Bolt Pricing.
v0: the Vercel plan decides, not the v0 plan
Vercel changed its terms in March 2026 with a clear split: "Hobby (including Trial Pro): Opted in for AI model training by default," paid "Pro: Opted out of AI model training by default," and "Enterprise: Opted out of any AI model training" - Vercel Changelog. The trap for v0 users is in the FAQ: "If you are a v0 Free or Premium user with a Vercel Hobby account, you will be opted into AI training by default." In other words, paying for v0 itself does not opt you out. What decides the default is the plan on your Vercel team, and paid Vercel Pro starts at $20 per month - Vercel Pricing.
Vercel's program also involves third parties. It says it shares "code and v0 prompt data, as well as telemetry on deployments and builds, errors encountered, and aggregate web traffic stats," which "may be used to improve our services and provided to third parties for AI model training purposes only." Vercel says this data is anonymized and redacted of "personal information, account details, environment variables, API keys, and other sensitive content" before sharing, which is a stronger safeguard than most builders document. The opt-out is in Team Settings (not Account Settings), then Data Preferences, with a per-project override under Project Settings, and like everyone else's it is not retroactive: "Data that may have been shared before you opted out can't be unshared retroactively."
One limit applies even after you opt out. Vercel states that "Pro and Hobby users cannot opt out of Vercel using customer data to improve and develop products... in ways that do not involve training an AI model." That is the "service improvement" category from Section 3 in practice: the training toggle switches off training, not every use.
Replit: no clear default, no documented switch
Replit is the hardest builder to pin down, and that is itself a finding. Its terms, last updated August 3, 2026, say "Replit reserves the right to access the content of your private apps for the purpose of troubleshooting, improving our service, and ensuring the safety and security of the Service" - Replit Terms of Service. Its privacy policy claims a "legitimate interest in using Personal Data for product development and internal analytics purposes, to improve the accuracy of our machine learning technologies such as code generation" - Replit Privacy Policy.
Neither page says in plain words whether project content trains models on consumer plans, and we could not find a documented setting to opt out of model training. The contrast with Replit's enterprise contract is sharp: there, Replit commits that it "will not use Customer Content to develop or improve Replit's products or services, train machine learning models, or create derivative works" - Replit Enterprise Agreement. When a vendor writes a precise no-training promise into its enterprise contract and a vague improvement clause into its consumer terms, the reasonable reading is that consumer content is not covered by the promise. If you use Replit for anything sensitive, ask support in writing which category your plan falls into, and keep the reply.
Base44: training on every plan below Enterprise
Base44, the builder Wix acquired, has the simplest policy in the category and the least protective. Its documentation says: "Enterprise: your data is not used to train AI models." And: "All other plans: your data can be used to train AI models" - Base44 Docs. We found no self-serve opt-out below Enterprise. That makes Base44 the only major builder in this guide where a paying, non-enterprise customer has no documented way out, which is why it sits at the bottom of the scoring table.
Base44's bluntness has one virtue: nobody can claim to have been surprised. For prototypes, throwaway experiments and public demos, that may be an acceptable trade. For anything you expect to become a business, the absence of an opt-out is a strong reason to keep the commercially sensitive part of the work elsewhere, which is part of the broader argument in our guide on when to graduate from a vibe-coding tool.
Figma Make, Google AI Studio, and what a toggle looks like
Designers building prototypes in Figma Make fall under Figma's content-training setting, which Figma introduced in August 2024. Its help center says content training "is turned on for Starter teams" by default, that only admins can change it, and that "if an admin turns off content training, new content and edits will not be used to train AI models" - Figma Help Center. Figma also commits not to let third-party model providers train on customer content. Whether that toggle covers Make output specifically is not spelled out on the page, so treat it as likely rather than certain.
Figma's toggle is a useful picture of what these controls look like across the industry: one switch, one sentence of explanation, and a link to a policy that does the real work.
Notice what the switch does not tell you: whether turning it off affects content already collected (for Figma, it does not), who else on the team can turn it back on (any admin), or what "de-identify" means for a design file. Every toggle in this guide hides the same three questions behind a single line of text. Google AI Studio, which many founders use to prototype apps on Gemini for free, makes the trade explicit instead: on unpaid services, "Google uses the content you submit... to provide, improve, and develop Google products," human reviewers may read it, and the terms say "Do not submit sensitive, confidential, or personal information to the Unpaid Services" - Gemini API Terms. Attach a billing account and those uses stop. That one sentence is the most honest summary of free AI tooling anywhere in this guide.
6. Coding Assistants: Copilot, Cursor, Claude Code, Codex and More
App builders are where most non-technical founders start, but the moment a product gets real, the code usually moves into an AI coding assistant or agent: an editor like Cursor, an agent like Claude Code or Codex, or GitHub Copilot inside an existing workflow. We compared those tools on autonomy and cost in our Claude Code vs Codex vs Devin guide. This section compares them on the question that guide did not cover: what they do with the code they see.
The coding-assistant market follows the same plan-boundary pattern as the app builders, with one important difference. Most of these tools sit on top of a consumer subscription (a ChatGPT plan, a Claude plan, a GitHub account) rather than a builder's own account. That means the training decision is often made by a setting in a product you think of as a chatbot, not by anything inside the coding tool. If you only remember one sentence from this section, make it this one: the toggle that governs your coding agent may live in your chat app's settings.
GitHub Copilot: default-on for individuals since April 24
GitHub announced on March 25, 2026 that "from April 24 onward we will begin using interaction data" from "Copilot Free, Pro, and Pro+ users to train and improve" its models - GitHub Changelog. Interaction data means inputs, outputs, code snippets and their surrounding context. GitHub's chief product officer was explicit that "Copilot Business and Copilot Enterprise users are not affected by this update" - GitHub Blog.
Two details matter for private code. GitHub says it does "not use private repository content at rest to train AI models," but the snippets Copilot sends while you work inside a private repository are interaction data, and those are in scope. And the license now extends to affiliates: "Affiliates may now use shared data for additional purposes, including developing and improving artificial intelligence," where affiliates include Microsoft. GitHub does say affiliates "do not include third-party AI model providers," and that its opt-out "applies to GitHub and its affiliates."
The opt-out is in your Copilot settings under "Privacy," and GitHub kept earlier choices: "If you previously opted out of the setting allowing GitHub to collect this data for product improvements, your preference has been retained." Like every other vendor, it works forward only: "Once you opt out, we stop collecting from that point forward." For organizations, the structural answer is Copilot Business at $19 per user per month - GitHub Docs, where training is excluded by contract rather than by a toggle someone might forget.
Cursor: Privacy Mode, and a new owner
Cursor's policy turns on a single setting called Privacy Mode. With it on, "Customer Data will not be used for training by Cursor," and Cursor "maintains zero data retention (ZDR) agreements with all providers." With it off, "we may use and store codebase data, prompts, editor actions, code snippets, and other code data and actions to improve our AI features and train our models" - Cursor Data Use. The setting is "available to anyone (free or Pro)" - Cursor Security, lives under Cursor Settings, General, and "for teams, Privacy Mode is enabled by default for all team members" - Cursor Help.
There are two exceptions worth knowing. Even with your own API key, "your requests will still go through our backend," and Cursor's help page states that "ZDR doesn't apply when you use your own API keys." And content that triggers abuse detectors "may be stored for investigation." If you are an individual on Free or Pro, open the setting and confirm it is on rather than assuming; Teams at $40 per user per month - Cursor Pricing is where it becomes the default for everyone.
The bigger change is ownership. On August 14, 2026, Cursor announced it "has officially been acquired by SpaceX," completing a process that began in April "when we announced our partnership with SpaceXAI to accelerate our model training efforts" - Cursor Blog. Two weeks later, CNBC reported that OpenAI planned to stop supplying models to Cursor, with a "proposed shutoff date" of November 12, 2026 - CNBC. Cursor's data-use page was updated on September 3 and now lists SpaceXAI among the providers whose retention terms apply. None of this changes what Privacy Mode promises, but it is a live example of the point in Section 2: the companies building your tools increasingly train models, and their ownership and model suppliers can change in a matter of weeks.
Claude Code: your Claude plan's setting decides
Anthropic changed its consumer terms in August 2025 so that "Claude Free, Pro, and Max plans, including when they use Claude Code from accounts associated with those plans" can be used for training if the user allows it - Anthropic. Users had until October 8, 2025 to choose, and the change applies "only to new or resumed chats and coding sessions." Claude Code's own documentation restates the rule: "We will train new models using data from Free, Pro, and Max accounts when this setting is on (including when you use Claude Code from these accounts)" - Claude Code Docs.
The retention consequence is stark and worth understanding before you decide. Users who allow training get a 5-year retention period; users who do not get 30 days. Feedback is handled separately: transcripts you share through the /feedback or /bug commands "are retained for 5 years." Under commercial terms (Team, Enterprise and API use, including API access through cloud platforms such as Amazon Bedrock), Anthropic "does not train generative models using code or prompts sent to Claude Code under commercial terms, unless the customer has chosen to provide their data."
To check where you stand, open Claude's settings, then Privacy, and look at the "Help improve Claude" toggle, which covers both the chat app and Claude Code on the same account. If you run Claude Code for long unattended sessions, as described in our guide to running Claude Code in auto mode, this setting matters more, not less, because those sessions read more of your codebase than any chat. On the Team plan, the pricing page states "No model training on your content by default," at $25 per seat per month billed monthly - Claude Pricing.
OpenAI Codex: one toggle, plus one you might miss
OpenAI is direct about personal plans: "When you use our services for individuals, such as ChatGPT and Codex, we may use your content to train our models" - OpenAI Help Center. The control is "Improve the model for everyone" under ChatGPT's Data controls, and OpenAI confirms that "if you use Codex on a personal ChatGPT plan, Improve the model for everyone also applies to your Codex tasks" - OpenAI Help Center.
There is a second Codex setting that the main toggle does not touch: "Codex has a separate Include environments setting... Changing your ChatGPT setting or opting out through the Privacy Portal does not change that setting." And there is a trap that catches careful users: when you give feedback (a thumbs up or down), "the entire conversation associated with that feedback may be used to improve our models, even if you've opted out." If you want your Codex work out of training, switch off both settings and avoid the feedback buttons on sensitive sessions.
By default, OpenAI does not use inputs or outputs from ChatGPT Business, Enterprise, Edu or the API to improve its models. ChatGPT Business costs $25 per user per month on monthly billing, with a minimum of two seats - OpenAI Help Center. A Codex CLI session signed in with a personal ChatGPT account follows the personal plan's setting, while one using an API key falls under the API's terms (Section 7).
Google: Code Assist's free tier is gone, Antigravity is the new question
Google's developer tools changed shape this year. "Starting June 18, 2026, Gemini Code Assist IDE extensions stopped serving requests for the Gemini Code Assist for individuals, Google AI Pro, and Google AI Ultra tiers," and the same notice applies to Gemini CLI usage on those tiers, pointing users to Antigravity instead - Google for Developers. For anyone still running Gemini CLI on a personal login, its documentation said prompts, answers and related code "may be used to improve Google's products, including for model training," with a single "Usage Statistics" setting as the control - Gemini CLI Docs.
Antigravity's terms are broad: Google uses your interactions "to evaluate, develop, and improve Google and Alphabet research, products, services and machine learning technologies," and "Google employees and contractors may access, view, review and use Interactions. If you don't want your Interactions used in this way, navigate to settings" - Antigravity Terms. Users on Google's own developer forum report that the only switch they could find was a telemetry toggle, and that they found no explicit training opt-out (a user report, not a Google statement) - Google AI Developers Forum. Until Google documents the setting clearly, treat a personal-account Antigravity session as training-eligible, and use it through a Workspace or Gemini Enterprise account for code you care about.
Amazon Kiro and JetBrains AI: clear switches, different defaults
Kiro, Amazon's agentic IDE, collects "content for service improvement from Kiro Free Tier users and Kiro individual subscribers" by default. The checkbox is under Settings, User, Application, Telemetry and Content, labelled "Content Collection for Service Improvement," and Free Tier inputs may be stored "for up to 60 days" for abuse detection - Kiro Docs. Enterprise users signed in through IAM Identity Center are excluded. Amazon Q Developer follows the same logic: Free tier content may be used "for model training," while "we do not use content from Amazon Q Developer Pro" - AWS Documentation.
JetBrains has the most founder-friendly default among the large vendors. Detailed code-related data, which "is used for product improvement and training JetBrains models," is shared only if you opt in on a paid license; "for licenses that are for non-commercial use, this setting is enabled by default." If you share, the data "is retained for up to one year" - JetBrains AI FAQ. That opt-in default on paid licenses is why JetBrains tops the scoring table: a paying customer who never opens a settings page is protected, which is the outcome every other row in the table makes you work for.
7. The Model Layer: APIs, Subscriptions and Retention
Section 3 introduced the two-layer structure: the builder, and the model provider behind it. Most of this guide is about the builder layer because that is where the 2026 changes happened. But a founder needs to know what the model layer promises too, because it is the floor every builder stands on, and because one popular setup (an AI coder running on your own subscription) puts you directly on the model provider's consumer terms.
The good news is that the commercial model APIs have converged on a protective default. The nuance is that "not used for training" does not mean "not stored," and the retention windows differ enough to matter when you are deciding which setup to trust with a codebase.
Commercial APIs: the quiet default that protects most builders
When a builder calls a model through a provider's API, the provider's commercial terms apply, and all three major labs exclude API traffic from training by default. OpenAI states that since March 1, 2023, "data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)," while abuse-monitoring logs are kept "for up to 30 days" - OpenAI API Docs. Anthropic says that "by default, we will not use your inputs or outputs from our commercial products (e.g. Claude for Work, Anthropic API, Claude Gov, etc.) to train our models" - Anthropic Privacy Center, and that API inputs and outputs are deleted "within 30 days of receipt or generation" - Anthropic Privacy Center.
Google's line falls between paid and unpaid use of the Gemini API: unpaid use can improve Google products with human review, while "when you use Paid Services... Google doesn't use your prompts... or responses to improve our products" - Gemini API Terms. This is why Lovable can truthfully say its agreements with model providers "restrict" their use of your content: the API layer was never the problem. The risk moved up a layer, into builders that now train models of their own.
When the AI coder runs on your own subscription
A growing number of founders run an AI coder such as Claude Code or Codex on their own consumer plan, either directly in a terminal (the setup in our guide to building a live app with Claude Code) or through a tool that drives those coders on their behalf. In that setup, the consumer terms apply: Claude's "Help improve Claude" setting and ChatGPT's "Improve the model for everyone" setting decide whether those sessions train models, not the API's commercial default.
Founden is one example of the second pattern, and since it publishes this guide its own setup should be stated plainly. Founden's terms grant it a license to customer content "solely as necessary to provide the Services to Customer" - Founden Terms, and Founden does not train models. Its desktop builds run Claude Code or Codex on the founder's own Claude or ChatGPT plan, so the training choice for those sessions sits with the founder's own account toggle rather than with Founden, and the build's prompts, steps and results still sync to and are stored by Founden. Its cloud builds run by default on a lab's model through Founden's own API account, which falls under the commercial API terms above. That is control over which setting governs your code, not immunity from anyone's policy, and it still leaves you one toggle to check on your AI plan.
Retention: training-on means years, training-off means weeks
The clearest way to see what the training decision costs is retention. Across the policies in this guide, allowing training extends how long your content is kept by an order of magnitude or more, because a dataset assembled for training is only useful if it persists.
The chart makes the trade concrete. On a Claude consumer plan, the same session is kept for 30 days if you decline training and for five years if you allow it, a factor of sixty. JetBrains keeps shared detailed data for up to a year. The commercial APIs sit at 30 days for abuse monitoring, and zero-retention arrangements exist for eligible enterprise customers: OpenAI's require "prior approval by OpenAI," and Cursor's Privacy Mode relies on ZDR agreements with its providers. The practical rule: if a vendor offers you a choice, declining training is also the single biggest lever you have on how long your code sits on someone else's servers.
8. The Price of Default Privacy
Every tool in this guide offers some route to keep your code out of training, but there are two very different kinds of route. One is a toggle on a cheap or free plan, which works only if every person on the team finds it and nobody turns it back on. The other is a plan where training is off by default and the exclusion is written into a contract, so nobody has to remember anything. The second is what a business actually wants, and it has a price.
Reading the plan boundaries as a price list is the first-principles move from Section 2: it shows what each vendor thinks your data is worth, measured by what it charges to give up the right to train on it. The chart below shows the cheapest monthly price, per seat on monthly billing, at which each tool switches training off by default.
Three things stand out. First, the spread is modest: $19 to $50 a month buys a contractual exclusion at every vendor in the chart, which is small next to what a leaked product strategy or a cloned feature set would cost. Second, the price is per seat on most tools, while Lovable prices by workspace credits and lets unlimited members join - Lovable Pricing, so a five-person team pays five times for most of these and once for Lovable, which changes the ranking for teams. Third, two vendors in this guide are missing from the chart for opposite reasons: JetBrains needs no upgrade because its paid licenses are opt-in only, while Base44 offers no self-serve tier with training off at all, only Enterprise.
Geography changes the arithmetic too. Bolt's default is already off for accounts in the EU, the UK and Switzerland on every plan, so a European founder gets for free what an American founder pays $30 a seat for. That asymmetry is not generosity; it reflects where data protection law makes default-on training risky, which is the subject of the next section. The way to apply all of this is simple: if a project is a throwaway experiment, use the toggle; if it is a business, budget for the plan, and treat the line item the way you treat paying for a password manager.
9. Beyond the Toggle: Law, Contracts and Acquisitions
A settings toggle is a promise a company makes in its own product, on its own terms, and can rewrite with an email. That is why it matters what sits underneath it: the law that constrains how a default can change, the contract you can sign to make an exclusion binding, and the ownership of the company that holds your data. None of these is a substitute for flipping the switch today, but together they decide how durable that choice is.
This section is not legal advice, and the rules differ by country. What it gives you is the map: which protections are real, which are weaker than they sound, and which ones a small founder can actually use without a lawyer on retainer.
The FTC's warning about quiet terms changes
In the United States, the most relevant guidance is a February 2024 post from the Federal Trade Commission's technology staff, titled "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive." It warns that it "may be unfair or deceptive for a company to adopt more permissive data practices," such as "using that data for AI training," and "to only inform consumers of this change through a surreptitious, retroactive amendment to its terms of service or privacy policy" - FTC. The post cites the 2004 Gateway Learning case, where a company changed its privacy policy to share consumer data "without notifying consumers or getting their consent."
Read with the 2026 changes in mind, the guidance explains a lot about how the vendors behaved. Lovable sent a dated notice five weeks ahead and said opting out before the effective date meant your data would never be used. Bolt published its new terms on September 14 but made them effective for existing accounts only from October 7. GitHub announced a month ahead and kept earlier opt-outs. Vendors are building their changes to be prospective and announced, which is exactly what the FTC said separates a lawful change from a deceptive one. The flip side is that a properly announced default change is probably lawful, so the protection the FTC offers is against surprise, not against the change itself. Your defence is reading the emails.
Europe: the right to object, and why Bolt excludes EU accounts
European founders have a stronger position. When a company relies on legitimate interest as its legal basis (as Lovable says it does), Article 21 of the GDPR gives you "the right to object, on grounds relating to his or her particular situation, at any time," after which the company "shall no longer process the personal data unless the controller demonstrates compelling legitimate grounds" - EUR-Lex. The European Data Protection Board's opinion on AI models adds that "whenever legitimate interest is relied upon as a legal basis by a controller, the right to object under Article 21 GDPR applies and should be ensured," and lists "an unconditional 'opt-out' from the outset" as a mitigating measure - EDPB.
Regulators have enforced this pattern on a much larger company. Before Meta began training on European users' public posts on May 27, 2025, Ireland's Data Protection Commission secured an "updated and easier to use Objection Form," in-app objection forms, and "access to the Objection Form for over a year" - Irish DPC. This is the background to Bolt's choice to keep EU, UK and Swiss accounts out of training entirely: default-on training in Europe invites regulatory attention that a US default does not. The limit to remember is that the GDPR protects personal data. Your source code and product strategy are often not personal data at all, so for a founder's commercial interest the toggle and the contract still do most of the work. We covered the wider European regime in our guide to making your AI app EU-compliant.
The EU AI Act adds transparency at the model layer rather than an opt-out. Its obligations for general-purpose AI providers have applied since 2 August 2025 - European Commission, and the Commission published a template to help providers "summarise the content used to train their model" - European Commission. That helps you learn what went into a model, not keep your project out of one.
California: disclosure, not a right to say no
California's AB 2013 requires developers of generative AI systems to post documentation about their training data "on or before January 1, 2026," including "a high-level summary of the datasets," whether they contain copyrighted material, whether they were "purchased or licensed," and whether they include personal information - California Legislative Information. It covers systems released since 2022, including free ones.
For a founder, AB 2013 is a research tool. If a builder trains its own models on user content and makes them available in California, its disclosures should tell you, at a high level, that user content is part of the training data. It does not give you a right to remove your project from a dataset. We found no US state law in force that grants an opt-out from AI training as such, so in the US the contractual and settings routes are the only practical ones.
Copyright will not rescue your code
Founders sometimes assume copyright would stop a company from training on their code. The courts are not moving in that direction. On September 16, 2026, the Ninth Circuit affirmed the dismissal of the DMCA claims in Doe v. GitHub, holding that Copilot and Codex do not "remove or alter" copyright management information "but instead create new works that never contained that information" - EFF. The remaining claims in that case are contract claims about open-source licenses, which is a narrow route that depends on license terms you control only for code you publish.
The money in AI copyright has flowed through settlements over pirated books rather than through rulings that training itself is unlawful: a federal court gave final approval in July 2026 to a $1.5 billion settlement in Bartz v. Anthropic, about $3,000 per book across more than 400,000 books, according to JURIST - JURIST. For a founder's private project, which you handed to the builder under its own terms of service, the lesson is that the terms you accept are the law that governs, which brings us to contracts.
Contracts: what a data processing agreement buys you
The plan boundaries in Section 8 are really contract boundaries. Lovable's Business and Enterprise workspaces are excluded "under your organization's Data Processing Agreement," and its privacy policy says "we set out this prohibition in our Data Processing Agreement." GitHub's Business and Enterprise exclusions are contractual. Replit writes its no-training promise into the Enterprise Agreement. The practical difference between a toggle and a contract is that a contract cannot be changed by a policy email: changing it takes your signature, or at minimum notice under the agreement's own terms. For a founder that turns the question from "did I remember the setting" into "what did we sign," which is a far more stable thing to rely on.
If you are signing up for a team or business plan, four clauses are worth checking before you rely on it:
- A no-training clause naming customer content explicitly, not just "personal data."
- Subprocessor terms saying the model providers the builder uses may not train on your content either.
- Retention and deletion periods, and whether deletion covers backups.
- Change of control language, so the commitments survive an acquisition.
The reason to read these rather than trust the plan name is that "business" plans differ more than their marketing suggests. Some put the no-training commitment in a signed agreement; others put it in a help-center page that can change like any other. The check takes ten minutes with the vendor's trust or legal page open, and if a clause is missing, a polite email asking for it in writing costs nothing and tells you a lot about the vendor.
Acquisitions are a data event
The last lever is one you do not control at all: who owns the company. The builder stack consolidated fast this summer. Cursor became part of SpaceX on August 14. Stripe announced on August 19 that it "has agreed to acquire OpenRouter, a leading AI model gateway and routing platform" - Stripe, which matters because a gateway sees every prompt routed through it. Lovable acquired Sutro on September 18. We mapped the earlier deals in our AI website builders market map.
An acquisition does not change a privacy policy on its own, and none of these companies announced data-policy changes with their deals. But privacy policies are written to let data move with the company. Lovable's is typical: in "a merger, acquisition, financing, or sale of assets, Personal Data may be disclosed to the other parties," and the policy keeps applying only "unless and until you are given notice of a different policy" - Lovable Privacy Policy. The new owner's incentives become the incentives that shape the next policy update. The defensive habit is cheap: when a tool you build with is acquired, re-read its data-use page within a month, and re-check your toggle when the next "updates to our terms" email arrives. If the commitment matters, it should be in a contract with change-of-control language, not in a setting.
10. Protecting What You Publish: AI Crawlers and Your Live App
Everything so far concerns the code and prompts you put into a builder. There is a second exposure that founders often forget: the app or website you publish. Once it is live, AI companies' crawlers can read it like anyone else, and some of them collect pages for training. This is a different problem with different tools, and it comes with a real trade-off, because the same crawlers that train models also feed the AI search answers that send you customers.
The control here is not a toggle in your builder but the instructions you give crawlers, either in a robots.txt file or at your network edge. The good news is that the major AI companies now separate their training crawlers from their search crawlers, so you can usually block training without disappearing from AI search.
Cloudflare's September 15 change
Cloudflare, which sits in front of a large share of the web, split AI traffic into three categories, "Search, Agent, and Training," and set new defaults on September 15, 2026: "For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default" - Cloudflare Blog. It also changed how multi-purpose crawlers are treated, saying crawlers such as Googlebot, Applebot and BingBot "will be blocked by customers who have selected to block Training." Some secondary coverage claims the new defaults also reach existing free-plan sites; Cloudflare's own post only commits to new domains, so check your settings rather than assuming.
The control panel Cloudflare published shows how the choice is now framed: three categories, each with a block, allow, or ads-only option.
The screenshot captures the important shift: "Training" is now its own row, separate from "Search." For a founder whose app is served through Cloudflare, this is the fastest way to stop training crawlers across a whole domain without touching code. Cloudflare's managed robots.txt also writes a Content Signals line such as Content-Signal: search=yes,ai-train=no, a machine-readable statement of preference whose specification says that allowing search "does not give permission to train AI models on your content" - Content Signals. Signals are a statement of preference, not a technical block, so they work best alongside crawler rules.
Blocking training crawlers in robots.txt
If you control your site's files, a few lines of robots.txt opt you out of the main training crawlers while leaving search alone. Each company documents its own token. OpenAI says "disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models," and controls its search crawler separately as OAI-SearchBot - OpenAI. Anthropic uses ClaudeBot and notes the rule must be set "for every subdomain" - Claude Help Center. Google's Google-Extended token covers Gemini training and "does not impact a site's inclusion in Google Search" - Google for Developers. Apple uses Applebot-Extended - Apple Support, and Common Crawl, whose archive feeds many training datasets, uses CCBot - Common Crawl.
A training-only block looks like this:
# Block AI training crawlers, keep search crawlers allowed
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
Most AI app builders let you add a robots.txt file to the project, or will add one if you ask the agent. Put it at the root of your domain, and repeat it on every subdomain you serve. Then decide deliberately whether this is what you want. If your growth plan depends on being cited by AI assistants, blocking the search and answer crawlers would hurt you, which is why the rules above block only training tokens. We went deep on that side of the trade-off in our guide to getting your site cited by ChatGPT and Claude. The sensible default for most founders is: allow search, block training on anything that is your product's substance, and revisit once a year.
11. The 20-Minute Opt-Out Audit
The previous sections explain the landscape. This one turns it into a procedure. The audit is designed for a founder or a small team and takes about twenty minutes the first time, most of it spent finding settings pages. It is worth doing today rather than when the product matters, for the one-way-door reason in Section 2: anything a tool trains on before you opt out stays trained.
The audit follows the two-layer structure from Section 3. You list every tool that touches your code, settle the builder layer for each, settle the model layer for any AI coder on a personal plan, then close the gaps that no toggle covers. The diagram shows the decision at each step.
The diagram's most important branch is the one people skip: the shared workspace. On Lovable, Bolt and the individual tiers of most tools, a single founder's opt-out does nothing for content a collaborator contributes, so the audit is not done until every person has confirmed their own setting.
Step 1: list every tool that sees your code
Start with an honest inventory, because the tools that matter are not only the ones you think of as "the builder." A typical founder's stack includes an app builder, a code editor or agent, a chat assistant used for debugging, perhaps a design tool, and the repository host. Each one that receives prompts, code or files is in scope, and each has its own policy.
Write the list down with the plan you are on for each, because the plan decides the default in almost every case. Include tools teammates and contractors use on your project, not just your own. A contractor fixing a bug with their personal Copilot Free account is sending your code under their account's settings, not yours, which is a conversation worth having once and putting in writing.
Step 2: switch off training, tool by tool
For each tool on a free or individual plan, find the control and switch it off. The table collects the settings covered in this guide, with the path each vendor documents. Labels change, so if your screen differs, search the vendor's help center for "model training."
| Tool | Where the setting lives | What it is called |
|---|---|---|
| Lovable | Account settings, Preferences, AI model training | "Use my Lovable content for model training" |
| Bolt | Avatar, Settings, General (turn off Forge first) | "Model training" |
| v0 / Vercel | Team Settings (not Account), Data Preferences | Opt-In / Opt-Out |
| GitHub Copilot | github.com/settings/copilot, Privacy section | The AI model training setting under "Privacy" |
| Cursor | Cursor Settings, General | "Privacy Mode" (turn on) |
| Claude Code | Claude settings, Privacy | "Help improve Claude" |
| Codex | ChatGPT Data controls, plus Codex settings | "Improve the model for everyone," "Include environments" |
| Kiro | Settings, User, Application, Telemetry and Content | "Content Collection for Service Improvement" |
| JetBrains AI | Settings, Appearance & Behavior, System Settings, Data Sharing | "Allow detailed data collection by JetBrains AI" |
| Figma | Team settings, AI (admins only) | "Content training" |
Two cautions apply across the table. First, an opt-out is not a deletion: it stops future collection, and anything already used stays used, which every vendor in this guide confirms in its own words. Second, feedback buttons are often a separate channel. OpenAI says feedback can make a conversation trainable "even if you've opted out," and Claude Code keeps /feedback transcripts for five years. On sensitive projects, do not rate responses.
Step 3: settle the team and the plan
If more than one person works on a project, decide whether you will police toggles or buy a plan that does it for you. Policing works for two people who talk daily. Beyond that, the cost of one missed setting usually exceeds the price difference in Section 8, and a workspace-level default is the only arrangement that survives new hires, contractors and forgetfulness.
When you move to a team plan, do the contract check from Section 9 at the same time: confirm the no-training language, the subprocessor terms and the retention period, and save a PDF of the terms on the day you sign. Policies change; the copy you saved is your record of what you agreed to.
Step 4: assume your prompts already leaked something
The audit's least comfortable step is the most important. Look back through your chat histories in each tool for anything that should never have been there: API keys, database passwords, customer emails, unreleased financials. Founders paste these in constantly, especially when debugging, and an opt-out does not remove them from logs, retention windows or datasets already assembled. GitGuardian found that commits assisted by Claude Code leaked secrets at about 3.2%, roughly twice the baseline, in its 2026 report - GitGuardian.
The fix is mechanical: rotate every credential that appears in a prompt, a chat or a committed file, and move secrets into the environment-variable settings your builder provides. Our pre-launch security checklist covers the full routine, and our guide to giving your AI agent an identity instead of an API key explains how to stop long-lived keys from ending up in prompts at all. Rotating a key takes minutes; it is the only step in this audit that actually undoes past exposure.
Step 5: block training crawlers, then set a recheck
Finish with the live app: add the robots.txt rules from Section 10, or set the Training category to blocked if your domain is on Cloudflare. Then put a recurring reminder in your calendar to recheck every tool's setting twice a year, and add a rule to your inbox that flags any email with "privacy policy" or "terms of service" in the subject from a tool you build with.
That last habit is the one that would have caught every change in this guide. Lovable, Bolt, GitHub and Vercel all announced their 2026 changes by email and on dated pages, weeks in advance. The founders who ended up opted in were not deceived; they were busy. A two-minute read when the email arrives is the cheapest privacy control there is.
12. What Opt-Outs Do Not Do
A guide that ends with "flip these switches and you are safe" would be doing the vendors' marketing for them. Opt-outs are worth doing, and the audit above is the minimum, but they have hard limits that follow directly from how training and data retention work. Knowing the limits is what tells you when a toggle is enough and when a project needs a different setup entirely.
The limits fall into four groups: time (opt-outs only work forward), scope (they switch off training but not every use), people (your content reaches tools through others), and drift (the meaning of a setting can change). Each one is documented in the vendors' own policies, which is the strongest evidence that they are structural rather than accidental.
Opt-outs only work forward
Every vendor in this guide says the same thing in different words. Lovable: opting out "does not retract content from training datasets assembled, or models trained, before you opted out." GitHub: "Once you opt out, we stop collecting from that point forward." Vercel: data shared before an opt-out "can't be unshared retroactively." Anthropic: your data "will still be included in model training that has already started." Bolt goes furthest, because its opt-out does not unwind "datasets already delivered to licensees," meaning copies may already sit with companies you have never heard of.
The structural reason is that a trained model is not a database you can delete a row from. Your project's influence is spread across billions of parameters, and no vendor offers a reliable way to extract it. That is why the timing of your opt-out matters more than its existence, and why the most valuable thing this guide can do for a Bolt user is to arrive before October 7 rather than after.
Training stops, other uses continue
The training toggle switches off one of the five activities from Section 3, sometimes two. Retention continues: 30 days on Claude without training, up to 60 days on Kiro's free tier. Abuse monitoring continues everywhere, and Cursor notes that content triggering abuse detectors "may be stored for investigation." Service improvement continues at Vercel, which says Pro and Hobby users "cannot opt out" of uses "that do not involve training an AI model." De-identified data continues at Lovable, which may use it "for any lawful purpose." And human review, where it exists, is governed by its own rules.
None of this is sinister, and most of it is necessary to run a service safely. But it means the honest promise of an opt-out is narrower than its label: your content will not be used to update a model's weights from now on. It will still be stored for a while, read by systems and sometimes by people, and used in aggregate to improve the product. If that residual exposure is unacceptable for a project, the answer is not a better toggle but a different place to do the work.
Your code reaches tools through other people
The most common leak is not a vendor's policy at all; it is a collaborator's settings. On per-account opt-outs (Lovable, Bolt, every personal plan), a co-founder or contractor who never opened their settings contributes content under their own default. Public projects are another route: anything you publish, from a public Lovable project to a public GitHub repository, can be read by training crawlers regardless of any builder setting. And the AI assistant a teammate uses for "quick questions" about your code is a tool in scope, whether or not it appears on your list.
The defence is partly technical (team plans with workspace-level defaults, private projects, robots.txt) and partly social: a short written rule, shared with everyone who touches the code, saying which tools are approved and that their training settings must be off. That one paragraph in your onboarding notes closes more gaps than any individual toggle.
Settings drift
A toggle you set in 2025 may not mean the same thing in 2026. GitHub's change is the clearest example: it carried forward earlier opt-outs, which was the user-friendly choice, but the setting it carried forward was originally about "product improvements" and now governs model training by GitHub and its affiliates. Lovable's control changed name between its August notice and its current documentation. Cursor changed owners. A setting is a moving target, and the only defence is the recheck habit from Step 5 of the audit.
When an opt-out is not enough
For some projects, the residual exposure of a cloud builder is too much even with every toggle off: regulated data, a genuinely novel algorithm, or work under a client's confidentiality agreement. The options then are structural. Commercial API access with zero data retention, where you qualify, keeps the frontier models while minimising what is stored. Running an open-weight model on your own hardware removes the third party entirely, at the cost of capability and setup effort, a trade-off we priced out in our guide to the best open-weight models to self-host. The broader case for keeping AI under your own control is in our guide to personal sovereign AI.
Most founders do not need to go that far, and pretending otherwise would be fear-selling. For a typical early product, the combination of switched-off toggles, a team plan once there is a team, rotated secrets and a robots.txt file reduces the exposure to a level that is reasonable for the stage. The point of this section is to know where that line is, so the decision is deliberate rather than accidental.
13. Where This Is Heading
The 2026 wave of default changes is not a coincidence of five companies independently deciding to be less private. It follows from the economics in Section 2, and those economics are getting stronger. Coding agents are improving fastest where they can learn from real building sessions, and the companies that host those sessions are the only ones who can collect them. As builders train more of their own models (Cursor's Composer, Lovable's reserved right to offer models through its AI Gateway), the value of the default to them rises with every model generation.
Reasoning from that, four developments look likely. They are inferences from incentives, not announcements, so treat them as hypotheses to test against the next policy email you receive.
Privacy becomes a priced feature. Today the price of default privacy is buried in plan boundaries (Section 8). The logical next step is pricing it openly: a discount for sharing data, or a surcharge for not sharing. JetBrains shows the opposite strategy, protecting paid users by default and competing on trust, and it is not obvious which approach wins. If the market rewards trust, expect more opt-in defaults on paid tiers; if it rewards price, expect more vendors to follow Lovable and Bolt down-market.
Builders become data vendors. Bolt's dataset-licensing clause is the first explicit statement from a major builder that user sessions are a sellable asset, and Vercel already shares anonymized data with third parties for training. If those datasets sell well, others will add similar clauses, and the opt-out will start to carry a visible cost to the vendor, which is exactly when defaults get stickier and settings get harder to find.
Geography splits the defaults. Bolt's EU exclusion and the DPC's work on Meta show that Europe's right to object makes default-on training costly there. The likely result is a two-speed market: off by default for European accounts, on by default elsewhere, with location decided by billing address and network location, as Bolt already does. Founders outside Europe will keep paying, in data or in plan upgrades, for what European founders get by default.
The decision moves to your account. As more founders run AI coders on their own subscriptions, the training decision shifts from the builder to the founder's account with the model provider. That is a gain in control (one setting you own, rather than one per builder), but it also concentrates responsibility: the consumer toggle on your Claude or ChatGPT plan becomes the most important privacy setting in your stack. The counter-scenario worth taking seriously is that consumer AI plans follow the builders and narrow the opt-out over time; if that happens, commercial API terms and team plans become the only durable protection.
14. Conclusion: A Decision Framework
The short answer to "does Lovable train on my data" is yes on Free and Pro since September 9, 2026, until you turn it off, and the turning off only works from that moment forward. The longer answer is that Lovable is one example of a shift across the whole category: GitHub in April, Vercel in March, Lovable in September, Bolt in October, each moving individual and free users to default-on training while keeping business customers out under contract. The reason is structural. Real building sessions are the training data that makes coding agents better, and only the builders can collect them.
That makes the useful question not "is this company good or bad" but "what does this project need." A weekend experiment and a company you intend to raise money for deserve different answers, and so does a solo project versus one with five collaborators. The same founder can reasonably accept a training default on one project and pay to remove it on another, as long as the choice is made on purpose rather than by not opening a settings page.
Three questions decide what you should do about it for any given project:
What is this project? A throwaway experiment needs only the toggle, or nothing at all. A product you intend to build a business on needs every toggle off and, once there is a team, a plan where training is off by default by contract.
Who touches it? If anyone besides you contributes, per-account opt-outs stop being reliable. Either confirm every person's setting in writing or move to a workspace-level default.
Where does the AI run? If an AI coder runs on your own Claude or ChatGPT plan, your account's privacy toggle governs those sessions, whatever the tool around it does. If a builder calls models through commercial APIs, the builder's own policy is the one to read.
Answering those three questions turns a confusing landscape into a short list of actions. Most founders will end up in the same place: run the 20-minute audit today, rotate any secret that ever appeared in a prompt, add a training-only block to robots.txt, budget $19 to $50 per seat for a default-off plan once the project matters and the team grows, and read every "updates to our privacy policy" email from a tool you build with. None of that requires a lawyer or a new stack, and all of it is cheaper before a training run than after one.
The deeper takeaway is about control. A toggle is a choice a vendor lets you make; a contract is a choice you hold; and a setup where the training decision sits on your own account, or on your own hardware, is a choice nobody else can quietly reverse. Founders do not need to treat every AI company as an adversary to take that seriously. They only need to notice that the default is a business decision made by someone else, and decide whether to accept it.
This guide reflects vendor policies as published on October 3, 2026. AI builders change their terms, defaults and setting names frequently, several of the policies here changed within the last six months, and Bolt's new terms take effect for existing accounts on October 7, 2026. Verify the current wording on each vendor's own page before relying on it.