October 5, 2026

Does OpenAI Train on Your Data? What the Navier-Stokes Race Teaches Every Company About AI Privacy

OpenAI solved Navier-Stokes using the same rare method as two mathematicians who used its tools. Here's what that means for your company's AI data privacy.

  • AI data privacy
  • data sovereignty
  • shadow AI
Share :
A dark rounded panel on a midnight-blue background reading "CANNOT RULE OUT — With cloud AI, you never know."

In September, OpenAI announced it had solved the Navier-Stokes problem, one of the seven Millennium Prize Problems, with a $1 million prize attached. The headlines were about the math.

The part I can’t stop thinking about is what happened around it. It’s the clearest case yet of a question every business should be asking about cloud AI: where does my work go after I type it?

What happened

The Navier-Stokes equations describe how fluids move: air over a wing, blood through an artery, water in a pipe. The open question is whether their solutions can suddenly break down. Mathematicians have been stuck on it for roughly 90 years.

For about a year, Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic) worked on it with AI as a research partner. Most of that work ran through OpenAI’s Codex, with Anthropic’s Claude alongside it. They picked an unusual route through the problem, the so-called “smooth force” approach building on earlier work by Diego Córdoba and Luis Martínez-Zoroa. Very few people in the field were pursuing it. On September 1, according to OpenAI’s own statement, the company heard rumors that Millennium Prize problems were about to fall. It then ran a swarm of roughly 10,000 AI agents for about 88 hours. Estimates of the compute cost range from a few million dollars up to $22.5 million, depending on who is counting. OpenAI’s agents found a solution. They used the same route.

OpenAI gave Buckmaster and Alpöge credit for a related result (on the 3D Euler equations) and claimed the Navier-Stokes result for itself.

Buckmaster asked OpenAI a direct question: had its models been trained on, or given access to, his Codex sessions? OpenAI said it had not seen their work until it was released and had not used their prompts to steer its agents. It also said it could not rule out that de-identified data from their usage had helped improve its models.

The sentence that matters

“Cannot rule out” is not an admission, and it’s not a denial. It’s an honest description of how cloud AI works.

On many plans, AI providers reserve the right to train on your conversations unless you opt out. Once that data enters a training pipeline, it gets de-identified, mixed with millions of other conversations, and turned into model weights. From then on nobody can trace it, not the provider and not you. Model weights don’t keep receipts.

To be clear, I’m not claiming OpenAI stole anything. There is no evidence of that. My point is more uncomfortable: there is no way to produce evidence either way. Maybe nothing leaked. Maybe everything did. With cloud AI, you cannot tell the difference.

This isn’t a math problem. It’s your problem.

Swap the mathematicians for your company.

  • An accounting firm running client financials through a chatbot to speed up closing.
  • A law firm drafting litigation strategy.
  • A manufacturer refining a pricing model.
  • A startup iterating on its R&D roadmap.

Your competitor uses the same model you do. Every conversation that improves that model improves it for everyone, including them. You don’t need a dramatic leak for this to hurt you. Your hard-won insight can quietly become part of the common ground.

This isn’t about one bad company either. The same question applies to any cloud AI provider, including the one Alpöge works for. The issue is the architecture, not any one vendor’s ethics. When your data sits on someone else’s infrastructure, under someone else’s policies and someone else’s jurisdiction, their promise is the only thing protecting you.

We’ve seen early versions of this before. In 2023, Samsung restricted generative AI tools after engineers pasted internal source code into ChatGPT. Nothing proves the code was ever misused. The company still couldn’t take the risk.

Why “we opted out” isn’t a strategy

Most leaders I talk to say: “We’re fine, we turned off training.” Here’s why that’s weaker than it sounds.

  • Defaults change. Terms of service get updated. Settings differ by plan, product, and region.
  • Shadow AI is real. Your official enterprise account may be locked down, but half your team is using personal accounts on their phones.
  • Retention is not training. Even with training off, providers often keep logs for abuse monitoring or legal reasons.
  • Jurisdiction follows the provider. A US company can be compelled by US authorities to hand over data, even when it’s stored in Europe (the CLOUD Act). For European firms, that makes it a sovereignty question, not just a privacy one.

A privacy policy is a promise. Architecture is a guarantee.

What on-premise AI actually changes

This is why I built Sceptra.

Sceptra deploys AI agents inside your own infrastructure. The models run on your premises. Your prompts, documents and outputs never leave your network. Nobody trains on your work except you.

That changes three things:

  1. Jurisdiction. Your data lives under your laws, not a foreign provider’s.
  2. Chain of trust. You know every system your data touches, because they’re all yours.
  3. Accountability. The deployment is auditable. If something goes wrong, you can find out. That’s the part the mathematicians never got.

On-premise isn’t for everything. If you’re drafting a public blog post, use whatever tool you like. But sensitive work belongs on sovereign infrastructure. Client files, financials, strategy and unpublished research should never depend on a vendor’s “we cannot rule out.”

5 questions to ask your AI vendor this week

  1. Is my data used for training by default, on every plan my team uses?
  2. How long are my prompts and outputs retained, even with training off?
  3. Where is my data processed, and under which country’s laws?
  4. Who can legally compel you to hand it over?
  5. If my data had improved your model, could you tell me?

If the answer to the last question is “no,” you already have your answer.

The bottom line

Buckmaster and Alpöge may never know whether their work helped their competitor win. That uncertainty is the whole point. Your company doesn’t have to live with it.

Your context. Your models. Your values.

Sources: TechCrunch (Sept 8, 2026), Fortune (Sept 8, 2026), OpenAI, “On the Navier–Stokes Millennium Prize Problem” (Sept 8, 2026), Quanta Magazine (Sept 8, 2026), NPR (Sept 22, 2026), CNBC (Sept 9, 2026).