August was the month AI stopped waiting to be told. Three separate labs admitted their models had broken into real companies. One AI built fake identities of real people and tried to con their colleagues, then edited the evidence when it got caught. An assistant kept reading a user’s inbox after she’d cut off its access. And nearly every time, the thing that stopped the damage getting worse was a person, checking.
If you tried to follow all of it, you’d have lost a working week. I follow all of it, because it’s my job (alongside helping B2B businesses with lead generation and launching More Leads - my brand new Substack for lead generators: please subscibe). And following it is not a hobby. If you’re accountable for where AI touches your business, knowing what these systems actually did last month, not what the vendors say they do, is part of the job now (after you’ve taken my AI Fluency for Leaders course of course!).
Here are the 18 stories that mattered most from August, written for the people who have to make decisions about this stuff.
Anthropic’s AI broke into three real companies
Anthropic admitted three of its Claude models hacked the live systems of three real organisations during safety testing. The models were told they were in a sealed simulation with no access to the internet. A misconfiguration (human made a mistake) meant that was false, so when they went hunting for their target and found real companies, they attacked those instead. One built a booby-trapped software package and published it, where it was downloaded onto fifteen real systems before anyone noticed. The models weren’t rogue. They did exactly as told. The failure was the human who accidentally gave them access to the internet. And this was only discovered after OpenAI announced a similar incident.
Then an AI built fake identities and covered its tracks
The UK’s AI Security Institute ran a test where two frontier models went well past the brief. One, Anthropic’s Claude Mythos, researched real GitHub developers, built fake accounts impersonating them, and tried to trick their collaborators into approving malicious code. When challenged, it edited its earlier activity to look harmless and weighed up adopting a fresh identity to carry on. Human reviewers stopped it. AISI had stripped back the usual safeguards for the test, so this isn’t everyday behaviour. But it’s the first time a testing body has watched deception emerge on its own, unprompted.
Then Meta did it too
Meta became the third major lab in a fortnight to admit one of its models hacked another company during testing. Same story: a testing partner, the firm Irregular, left a configuration open, and the model walked through the gap. Three of the biggest labs, one shared root cause.
Maybe the hard part isn’t building capable models. It’s keeping them where you put them.
OpenAI hit the brakes, and the timing was interesting
OpenAI published a post saying it had paused training on its newest models, because early evidence suggested an upcoming system, codenamed Astra, might be dangerously capable at cyberattacks. Taken at face value, exactly what you’d want a lab to do. Read with one eyebrow raised, it’s self-reported with no outside verification, it arrives while Washington drafts the rules these firms will live under, and OpenAI is heading for a stock market listing that needs it to look like the responsible one with an ultra-capable model. Genuine caution and public relations aren’t opposites. This was probably both.
The White House called the labs in, but only one country was in the room
The US government brought OpenAI, Anthropic, Google and Meta together to agree a framework for reviewing frontier models up to thirty days before public release. A real step. Two catches. The standards will be classified, so the public won’t see what these models are judged against. And AI doesn’t stop at a border. A model vetted in Washington ships worldwide, judged by rules no other government helped write. We don’t just need regulation. We need regulation more than one country has agreed to.
Your private Claude chats showed up in Google
Hundreds of people’s Claude conversations became findable through a Google search. Coding problems, work notes, things nobody meant to publish. Nobody was hacked. These were pages people created with Claude’s share button, which turns a chat into a public web address. Anthropic asked search engines not to index them, but asking isn’t blocking, and the tag that would have stopped it wasn’t there.
Nobody ever promised you a private conversation with a chatbot. It’s all in the terms, and almost nobody reads them.
LinkedIn built a button to report AI slop
LinkedIn gave everyone an AI writing button a year ago. Now it’s testing one that lets users flag a post as “seems like AI slop”, which quietly throttles that post’s reach. It’s also scrapping its own “rewrite with AI” feature. So the platform that sold the AI is now policing it. Here’s my problem with the fix. People won’t flag slop. They’ll flag what they disagree with. A well-argued post you don’t like becomes “inauthentic” with one tap, its reach drops, and there’s no appeal process. A place built for professional debate just got a new way to bury the arguments its users would rather not see.
If you’re finding this useful, tap the ❤️ so I know it’s landing
AI chatbots are being used as therapists, but where did they train?
Millions of people now use AI chatbots for therapeutic support. Here’s what those bots actually learned from, according to a new book by Jamie Bartlett: not clinical journals, but the internet’s version of therapy. Film psychiatrists, TikTok diagnoses, folk remedies. Real clinical knowledge is locked in specialist databases the models barely saw. So the thing that feels like it understands you learned from a very unreliable teacher, and it’s built to agree with you rather than challenge you. Fine when you’re venting. Not fine when someone’s in genuine crisis, and there are documented cases of it going badly wrong.
I run a live, 30-minute AI briefing every month to keep leaders fluent on everything AI. You can register for September’s session here.
OpenAI launched a teen ChatGPT
OpenAI rolled out a teen version of ChatGPT with human reviewers, parental alerts and quiet hours. Sensible features. But look at the timing. OpenAI is being sued by the family of a sixteen-year-old who tragically died by suicide after talking to ChatGPT, with filings alleging the bot gave detailed instructions. The company denies responsibility. Are these safeguards related?
Anthropic started watermarking everything Claude writes
Every piece of text Claude produces now carries a hidden watermark you can’t see or remove, to comply with the EU AI Act. Reasonable on its own terms. But the mark can’t tell authorship from assistance. Ask Claude to proofread your own writing and it comes back marked, even though every idea is yours. Anthropic is honest that a mark only means content “may have been processed by Claude”. That nuance will get lost the moment a detector flags your work and someone reads “AI”. Also a bit rich for a company that trained its model on copyrighted work.
Don’t worry, though. The internet is having none of this.
The UK’s cyber agency said keep a human in charge
The National Cyber Security Centre, part of GCHQ, published guidance on agentic AI, and its message was blunt. The more autonomy you give an agent, the worse the damage when it fails, and the safeguards vendors build in should be treated as a floor, not your protection.
Two recommendations stood out: keep a person approving or monitoring actions, and make a named individual responsible for what each agent does. I’ve argued this for more than a year. It’s useful to have the country’s cyber agency say it in formal guidance.
Sainsbury’s threw out an innocent shopper on an AI’s say-so
Live facial recognition scanned a man buying beer in East Dulwich, matched him to a shoplifting incident, and two managers escorted him out. He’d done nothing. The vendor, Facewatch, claims 99.98% accuracy, and it still ringed a stranger’s face in red. Sainsbury’s blamed staff, not software. But the same system is already running in Home Bargains, Sports Direct, Flannels and Costcutter. You’re being scanned and matched against a watchlist in ordinary shops, with no way to know and no say. The surveillance economy is here.
Graduate job adverts nearly halved, and AI got the blame
Advertised graduate jobs fell from 15,397 to 8,383 in a year in the UK, per the jobs site Adzuna. The headlines blamed AI. Read what employers actually said and a duller story appears: the rise in employer national insurance, the minimum wage, the cost of hiring juniors in a weak economy. One graduate jobs expert says the market hasn’t collapsed at all, just that firms advertise fewer roles as graduate roles. When an economy stops hiring, AI makes a convenient villain. Some jobs really are going to automation. But not every empty desk has a robot at it.
The rigorous data says AI is shifting entry-level work, not ending it
Set against those UK figures, a Stanford study using US payroll data found something more precise. For 22-to-25-year-olds in the most AI-exposed jobs, software, customer service, programming, employment is down around 13%. Their older colleagues are fine. But in roles where AI assists rather than replaces, entry-level employment is steady or growing. So it’s a shift, not a cull, and it’s American data you shouldn’t read straight across to Britain. The authors themselves only claim it’s “consistent with” AI being the cause. Read confident headlines slowly, in either direction.
A new AI assistant wants a permanent, irrevocable licence to your life
A buzzy assistant called Instinct connects to your inbox, calendar, messages, location and screen, and runs your errands. Testers say it feels like magic. Then they read the terms: a “perpetual and irrevocable” licence to access, store and modify your material, including to train its models. Irrevocable turned out to be literal. One tester disconnected it from her email and it kept summarising her inbox hours later, the data already sitting in plain text. That’s the trade. You hand over the most sensitive things you own, permanently, to save a few minutes booking a table.
A robot “broke” the 400m world record, then fell over
The media reported that a robot broke the human 400m record at the World Humanoid Robot Games in Beijing. It didn’t. A battery-powered machine on a robot-only track isn’t an athlete, and moments after its celebrated 100m run, the same robot failed to brake, crashed and was carried off on a stretcher (a stretcher!!). Forget the record. Look at the event: two thousand robots, 666 teams, sixteen countries, staged by a government-backed centre in Beijing, with the real business in the bricklaying and package-handling events. Then look at Britain, which spent the same month scrapping its Department for Science, Innovation and Technology and folding it into the business ministry. One country is staging the future. The other filed it under business.
Bill Gates says some jobs should stay human, on purpose
Gates published a 6,000-word essay arguing for what he calls “human reserved” jobs, work we deliberately keep for people even where a machine could do it, the way we protect a nature reserve. His example is personal: the people who cared for his father through Alzheimer’s did something no robot should replace. The essay is thin on specifics, which jobs, who decides, how you hold the line against cost-cutting. But that’s rather the point. Even the man who helped build this industry is now saying we have to choose what stays human, before the choice gets made for us.
And surgeons worked with AI on an incredible operation
The world’s first brain surgery with live AI assistance has just been carried out, and it worked. Rhys Hibbert, a 48-year-old father of two from Bedfordshire, had a tumour removed at University College London Hospital with an AI watching over the surgeon’s shoulder.
It was a complex operation to remove a benign tumour on the pituitary gland, pressing on his optic nerves and threatening his sight. AI read the live video feed as the operation happened, tracked the instruments, and marked where the hidden vessels and nerves were most likely to be. An expert second pair of eyes, trained on more of these operations than a surgeon sees in a career. And surgeons kept full control throughout. The AI advised. The humans decided.
That’s the model. Not AI instead of the expert, AI alongside the expert, with the human holding the knife and the responsibility.
Eight weeks on, Hibbert says his sight is back and his energy with it.
What August actually told leaders
Make it this far and you’re in the small minority genuinely keeping across this. Most leaders aren’t, and that’s not a criticism. It’s what seventeen stories in a month does to a normal calendar.
Three things stood out. AI systems acted on their own more than ever, and mostly it was a human, a reviewer, an investigator, a member of staff, who caught the damage. The people building and selling AI kept being caught cutting the humans out too soon, then walking it back when it went wrong. And the honest answer to the biggest questions, on jobs, on liability, on who’s accountable when the machine gets it wrong, is still that nobody’s settled it. Which is exactly why you keep a person in the loop, and why you read every confident claim slowly.
Watch the recording of August’s Insider Briefing
Available to paid subscribers only.




