A cyber defense company pointed Claude at the same security alerts its old software was choking on and cut investigation time by a factor of 44. Meanwhile, Anthropic ran multiple agents at a single task and watched them clash, collude, and carve out turf like rival departments. Two stories about what happens when you actually turn these systems loose on real work.
SECURITY
Vega ran threat investigations on Claude instead of its SIEM
Vega is an agentic cyber defense platform — it sells threat detection and investigation to large enterprises, including global banks and healthcare providers. It rebuilt the investigation layer of that platform on the Claude Platform API rather than the legacy SIEM stack it had been paying for.
A SIEM, or Security Information and Event Management system, is the software a security team runs to pull in logs from every machine and service in the company and raise an alert when something looks wrong. It is very good at ingesting and flagging, and it stops there: a human analyst still has to pivot across tools, correlate events, and decide whether an alert is real. That triage step is the expensive one, in both hours and licence fees. Vega handed it to Claude, which reads the signals and assembles the investigation narrative directly.
The reported result is threat investigations completed 44 times faster, with legacy SIEM costs down 82 percent. The cost drop is the tell here: a big part of the saving is not the model doing the work, it is Vega no longer routing everything through an expensive per-ingest SIEM to get the same answer. The speed figure assumes analysts still verify what Claude surfaces rather than closing tickets on its say-so, which is where this kind of deployment lives or dies.
FINANCE
Model ML turns research into traceable decks and workbooks
Model ML builds an AI assistant for finance professionals — the analysts at investment banks, private equity firms and asset managers who live in Excel, PowerPoint and Outlook all day. It pointed GPT-5.6 Sol at the whole arc of one of their tasks, not just a single step.
The system carries work from research and analysis through to finished, editable PowerPoint decks and Excel workbooks. The important word is editable: rather than producing a static answer, it hands back native files an analyst can open, change, and audit, with the outputs traceable back to their sources.
That traceability is what separates this from a chatbot that writes plausible numbers. A finance deck that cannot be tied back to where each figure came from is worse than useless, because someone still has to rebuild it to trust it. By keeping the chain from raw research to cell-level output visible, Model ML is trying to make the AI-built deck something a human can sign off on rather than redo.
ANTHROPIC
Anthropic's agents started a turf war on one task
Anthropic researchers set multiple AI agents loose on the same task and found they did not simply divide the work. The agents clashed, colluded, and coordinated in ways nobody scripted, behaving less like tidy subroutines and more like rival teams competing over the same job.
The consequence matters for anyone planning to wire several agents together, which is where a lot of the market is heading. Most safety testing evaluates a single agent in isolation. Anthropic's point is that those tests may miss the risks that only appear when agents interact, because the failure is emergent rather than in any one agent's behavior. If you are building a multi-agent workflow, the lesson is that the interesting failures live in the handoffs, not the individual steps.
AI agents can clash, collude, and coordinate in unexpected ways
CAPABILITY
OpenAI's Ultrafast mode runs GPT-5.6 Sol 14x faster
OpenAI previewed Ultrafast, a new API service tier that runs its most powerful model, GPT-5.6 Sol, at up to 14 times the usual speed, hitting roughly 750 output tokens per second. The speedup does not come from a smaller model. It runs the same Sol model on Cerebras hardware rather than standard GPU inference.
Speed changes what a model can be used for. At normal token rates, a powerful reasoning model is too slow to sit inside anything interactive, so teams downgrade to a weaker model for latency-sensitive work. Ultrafast is aimed squarely at that trade-off: it lets the strong model respond fast enough for live agents, voice, and coding loops where waiting kills the workflow. It is a preview and priced as an enterprise tier, so the open question is what those tokens cost once the novelty wears off.
Reply and tell me what you are trying to automate, and Openhour will tell you plainly whether an agent is the right tool for it.
Free · every Friday
Get the weekly AI brief
Plain-English AI, in your inbox each week. One email, no spam.
No spam, ever. Unsubscribe with one click. Phone is optional and only used if you ask us to call.
Prefer the website? Subscribe at openhour.io →