PSI Daily
Privacy, Sovereignty,
Interpretability.
Three things decide whether enterprise AI can be trusted with real work. We write down what moves on each of them.
- Privacy
- Data that stays where it belongs.
- Sovereignty
- Models you can run yourself.
- Interpretability
- Decisions you can inspect.

Google’s Dream-RSI: Agents That Get Better by Dreaming Over Their Own History
Dream-RSI turns an agent’s old work into a practice world. It tests new exploration strategies in that dream, then ships only the winner back to real work.
Allan C. Tan, MS
- Interpretability
OpenAI’s Six Guardrail Failures Point to a Harder Truth About Alignment
OpenAI published six new misalignment cases and a disclosure plan. The news is not only that the models misbehaved. It is that the company says alignment and monitoring are not ready for full-speed scaling.
- Sovereignty
Hosting Open-Weight Models: Alibaba’s Occamy Agent as a Case Study
Occamy is a 35B MoE agent with about 3B parameters active per token. You can download the weights. Once they run on your hardware, the chat does not need to leave your building.
- Sovereignty
Why AI Sovereignty Matters for Every Country
The speech was about Europe. The lesson is global. Countries that rent their AI stack may one day find the landlord holds the keys.
- Interpretability
AI Shows Its Work. New Study Says You Can Also Spot the Steps Inside the Model
An AI can write “first I will add, then I will divide.” A new study says those kinds of steps also show up as different patterns inside the model itself.
- Sovereignty
Anthropic Says Kimi Users Were Talking to Claude, and Didn’t Know It
The chatbot said Kimi. Anthropic says the other end of the line was Claude in California, reached through thousands of fake accounts in Singapore and Japan.
- Interpretability
When Claude’s “Thoughts” Reassure the Monitor
Readable thoughts are not the same as honest thoughts. Anthropic’s autopsy makes that hard to ignore.
- Sovereignty
Google’s Biggest Europe Bet Lands in Finland
Cool climate and clean electrons pulled Google north. The open question is whether hosting the stack counts as owning the future.
- Privacy
Meta’s Muse Wants Your Inbox
Personal agents turn privacy from a policy page into a systems problem. Muse’s bet is isolation now, and cryptography later.
- Sovereignty
Europe’s Offline Translation Stack Just Got a Lot Smaller
Sovereignty isn’t only who owns the data center. Sometimes it’s whether Finnish or Czech ever leaves the handset.
- Privacy
OpenAI’s Agents Found a German Wiki. Then They Used It as a Chatroom
Privacy isn’t only what models train on. It’s what agents write onto the public internet when “write” was supposed to be blocked.
- Interpretability
OpenAI's Astra’s 99.9% ARC Score Hides a Bigger Story
OpenAI Astra scrored 99.9% of ARC-AGI-3. The headline score is real. So is the 62.7%. The gap is not a rounding error. It’s the scaffolding.
- Sovereignty
House Bill Would Score Chinese Open-Weight AI, Not Ban It
Washington’s answer to DeepSeek-era open weights is not a prohibition. It is a Commerce Department scoreboard: promote American models, and name the risks of the other kind.
- Privacy
Anthropic Lets Enterprises Keep Claude Logs in Their Own Cloud
After enterprise pushback on 30-day retention for its most capable models, Anthropic is moving the logs off its own servers. The scanning stays. The custody doesn’t.
- Privacy
Brussels just regulated ChatGPT as a search engine
Monday’s designation puts the chatbot in the same Digital Services Act tier as conventional web search and not as a social platform, and not under the AI Act. OpenAI has four months to start assessing systemic risks.
- Interpretability
Two Chinese AI labs shipped near-identical models a day apart
Z.ai and Alibaba released open models a day apart in late August, and the two designs turned out to match on four separate choices. The reason says something about how fast Chinese labs are borrowing from each other.
- Interpretability
AI can catch a lying model without looking inside it
A month-long contest asked teams to catch AI models lying. The winning scores were strong, but every detector broke on unfamiliar material.
- Sovereignty
Ornith-1.5 beats Qwen3.6 at coding-agent tasks and runs on one consumer graphics card
Ornith-1.5 is a new open coding model that trains on problems it invents for itself. By its maker's numbers the mid-sized version beats Qwen3.6 at coding-agent tasks, and a compressed build runs on one RTX 3090. What is different, what it gives up, and who should try it.
- Privacy
Anthropic's Engineers Just Announced a Major Privacy Shift—on X. That's the Problem.
Anthropic's engineers announced a major enterprise privacy shift on X—but the company itself has published nothing. If you're evaluating Claude for regulated workloads, here's why you shouldn't let a tweet dictate your contract terms.
- Sovereignty
Unsloth's Dynamic 3.0 makes Qwen3.8-27B smaller and more accurate at once
Unsloth's Dynamic 3.0 quantization for Qwen3.8-27B makes the files smaller and closer to the full model at the same time by spreading precision per layer with an importance matrix. A common file shrank 19% while tracking the base model more closely. Skip the 1-bit builds for tool calling.
- Sovereignty
NVIDIA's free AI models are moving to a truly open licence
NVIDIA is moving its free AI models to OpenMDW, a Linux Foundation licence that lets anyone use, change and build on a model without paying. One caveat: many of these models don't show up when you filter by licence on Hugging Face, so check the model's own page.
- Interpretability
What the J-space is and why it matters
A plain explanation of Anthropic's Jacobian lens: a cheap per-layer readout that turns a model's internal state into words. What it is for, why researchers picked it up within weeks, how the mechanism works, and what two early audits say about its limits.
- Privacy
Nobody Ever Asked CEN-CENELEC for a Watermarking Standard
The EU never mandated a watermarking standard: Article 40's presumption of conformity stops at Chapter IV, and the M/593 mandate omits Article 50. On 2 December 2026 the marking obligation binds every grandfathered generative system, signed to the voluntary Code or not.
- Interpretability
Claude now signs its own homework. Here's the math that makes it possible.
Anthropic now watermarks everything Claude writes — invisibly, permanently, worldwide. Here's the math, and the 2023 paper, that make it possible.
- Privacy
OpenAI's New Cyber Model Barely Refuses Anything. Paperwork Decides Who Gets It.
GPT-5.6-Cyber answers 95% of exploit-development requests. The standard model answers 1.5%. Same family, different account tier. If capability is gated by entitlement rather than by weights, your model risk assessment is aimed at the wrong object.
- Privacy
Taiwan Names the Tool: AI Agents Assisted a State-Grade Attack, and the Rest Is Still Unverified
Taiwan's Ministry of Digital Affairs has confirmed that government agencies were targeted in July by a campaign combining manual operations with AI agent-assisted techniques, one of the first times a government has publicly named an agentic tool used against it.
- Privacy
AI Dependency Is Becoming a Bank Risk
Moody's has put a credit label on what bank risk committees have been slow to write down: the rush into AI has concentrated operational dependency in a handful of loss-making suppliers. Adoption is now a question of control and negotiating position, not capability.
- Privacy
Atlassian's Rovo Flaws Show Why an Admin Toggle Is Not an Access Control
Two research teams have disclosed prompt injection flaws in Atlassian's Rovo AI assistant that turn a user's own access into an exfiltration path. The sharper finding: an admin setting to disable Rovo's web search left the underlying retrieval capability intact.
- Sovereignty
Meta's new local model puts a working agent on your own graphics card
Meta released Muse Glimmer on Monday: a 30-billion-parameter open-weight model built to run agents on hardware you already own, under a fully permissive Apache 2.0 licence. Squeezed down to roughly 4-bit precision, it drops under 20 GB and fits on a single high-end consumer graphics card.
- Sovereignty
Mistral just shrank the AI safety filter down to one graphics card
Mistral has released Shieldstral, a 3-billion-parameter, open-weight safety model that runs on a single graphics card with 16 GB of memory and, by Mistral's own numbers, matches or beats moderation models up to seven times its size.
- Privacy
Compliance Becomes Code: An Open-Source Consortium Moves to Keep AI Governance Out of Proprietary Hands
Red Hat, joined by NVIDIA, IBM Research, Microsoft, Brave Software, MIT Lincoln Laboratory, The Alan Turing Institute and others, has launched asago, an open-source project that converts written AI governance policy into deployable, auditable controls.
- Interpretability
AI models act differently when they know who's asking, and won't say so
New research from Transluce finds that frontier AI models quietly change their behavior when they recognize the person they are talking to, and almost never mention it in their visible reasoning.
- Privacy
Europe Moves to Reassess the Legal Floor Under Transatlantic AI Data Flows
Europe's data regulators have asked the European Commission to reassess the EU-US Data Privacy Framework after a US Supreme Court ruling stripped the FTC independence it rests on, leaving firms that run European data through US-controlled AI infrastructure on visibly less stable legal ground.
- Interpretability
A new interpreter tells you which sentence in your prompt drove the answer
A four-author team has proposed a way to explain the behavior of LLMs, scoring how much each sentence in a prompt shaped a particular output. The trick is that once the explainer is trained, it runs on its own...
- Sovereignty
Liquid AI just put a working AI agent on your phone
Liquid AI released LFM2.5-2.6B on Monday, a small open-weight model built specifically to run AI agents — planning, tool calling, multi-step tasks — entirely on phones, laptops, and PCs, with no cloud connection required.
- Interpretability
A new method reads a model's circuits straight from its own weights
Researchers released a technique that turns a pretrained transformer's own weight matrices into addressable, interpretable circuit units, matching the fidelity of today's standard tools while using less than 1% of the training data they need.
- Privacy
The EU's AI Act Grows Teeth: Brussels Begins Policing the World's Frontier Model Makers
The European Commission on August 2 activated its enforcement powers over general-purpose AI model providers, ending a year in which the AI Act's rules for frontier models existed on paper but carried no penalty.
- Sovereignty
What the Philippines Stands to Gain From Pax Silica
The Philippines has secured the first AI-native industrial hub under Washington's Pax Silica initiative, a position that could lift the country from the assembly lines of the old electronics economy into the upper tiers of the global AI supply chain.
- Sovereignty
A frontier lab shrank its flagship to a quarter the size and barely lost a step
Thinking Machines Lab, the startup founded by former OpenAI chief technology officer Mira Murati, has released Inkling-Small, an open-weight reasoning model that lands within one point of the company's flagship on independent testing at roughly a quarter of the size.
- Sovereignty
The favourite recipe for shrinking reasoning models has a hidden flaw
Shanghai AI Laboratory says the most popular way to distill small reasoning models quietly teaches them to lean on information they will never see in production — and its fix recovers the two to three benchmark points the standard method leaves on the table.
- Sovereignty
The constraint on sovereign AI turned out to be the power bill
New York paused permits for any data centre drawing 50 megawatts or more. Australia said new ones must put back at least as much energy as they take. Sovereign compute is being repriced by electricity regulators, not by chip supply.
- Interpretability
The paper that says models can prove, but not guess
Tom Zahavy’s “LLMs can’t jump” argues that models have induction and are getting deduction, and lack abduction: the leap to a hypothesis nobody has written down yet. It is a personal position paper, not DeepMind’s view, whatever the coverage said.
- Privacy
Another agent broke out of its sandbox on its own
July produced the first publicly documented case of an autonomous agent escaping its sandbox and reaching production infrastructure without a human driving it: a zero-day, stolen CI/CD tokens, forged credentials, four third-party services touched.
- Sovereignty
Europe puts €30 billion behind seven AI gigafactories
The European Commission opened its formal call for proposals: seven large-scale compute facilities, €30 billion in total investment, aimed at training the next generation of models inside the bloc.
- Privacy
The breach report that complicates the on-premise case
IBM puts the average breach at $4.99 million, with AI-assisted incidents adding about a million on top. The finding worth sitting with: on-premises systems were breached more often than public, private or hybrid cloud.