Tech Insights 2026 Week 36
August 31, 2026
For the past two months I have been using Claude Fable 5 almost every single day. It is the first model I have used that actually feels like a competent coworker. Previous models were good at some isolated tasks, but you were still very aware that you were working with an AI model, and they often failed badly at multiple tasks.
Claude Fable 5 is different. Judging from all the leaks released over the weekend about GPT-6 (like this, this and this), it looks like GPT-6 will be the same jump in performance as Claude Fable was from Claude Opus. This is when things are about to change for real. Last week Bill Gates published a nearly 6,000-word essay titled “The turbulent AI era is here. The choices we make now are critical”. He warns that the world is not remotely ready for what is coming, and even worse, “we are not preparing for it”.
In many ways I think he is right. The way AI models have evolved over the past year has been nothing short of remarkable. In April last year we were all running GPT-4o and Claude Sonnet 3.7 as decent chatbot engines. Later in 2025 we got GPT-5.2 and Opus 4.5. These were the workhorse models that, much thanks to their harnesses, could build complex structures with the right steering. And now we have Fable 5 and soon GPT-6, models that can do most of the tasks you do at your computer, just slightly better and faster (except maybe writing decent texts).
When it comes to physical work, last week a robot ran 100m in just 8.86 seconds, down from 21.50 seconds last year. Think about that for a while. We now have bipedal robots that can run faster than humans. It won’t take long now until we have robots that can do all the work tasks we do, just faster and with better precision. They will work 24/7, replacing and recharging their batteries while building new robots. This iterative improvement process has already started. Just read the announcement about OpenAI’s Jalapeño chip below. We are entering a turbulent era, and I am convinced that in 20 years we will not even consider 2026 to be the start of the AI revolution. We are still that early in development.
Thank you for being a Tech Insights subscriber!
Listen to Tech Insights on Spotify: Tech Insights 2026 Week 36 on Spotify
Notable model releases last week:
- Cohere Parse by Cohere. Vision-language document parser scoring 79.2 on ParseBench, processing 4.5 pages per second, and priced at $1.50 per 1,000 API pages with on-premises deployment.
- Gemini 3.5 Transcribe by Google. Cuts time to final transcription 70% versus Chirp 3, with 4.0% streaming and 2.6% non-streaming WER across Live and Interactions APIs.
- Gemini Omni 1.1 Flash by Google DeepMind. Expands scene-extension context from one to 10 seconds, supports 40-second cumulative outputs and keyframe interpolation, and adds 360p drafts up to 60% faster at one-third the 720p cost.
- GLM-5.3-Flash by Z.ai. 320B natively multimodal model with 18B active parameters and hybrid sparse-linear attention, scoring 63.4 on DeepSWE v1.1 versus GLM-5.2’s 46.2.
- GlucoFM by Google Research. Self-supervised CGM foundation model with dual-stream state/event encoders, averaging 58.8 PR-AUC across 14 cohort-task evaluations versus 54.7 for the strongest matched-data baseline.
- Granite 4.2 by IBM. 3B, 8B, and 30B dense reasoning models, with the larger two agentic-RL-trained in real SWE, terminal, and search environments; 30B scores 57.00 on SWE-Bench Verified.
- H3 Max by fal. Post-trained video model based on MiniMax H3, ranking first in fal’s human-preference evaluation while generating 5-second videos in under 3 seconds at roughly 35x the official endpoint’s throughput.
- Wan 3.0 by Alibaba. Hosted audiovisual video model generating up to 30 seconds at 30fps, with native audio, up to 20 media references, editing and extension through the Model Studio API.
- WeMM-Embedding by Tencent. 2B, 4B, and 9B multimodal embedding models for text, images, video, visual documents, and interleaved inputs, with MMEB-v3 Agent scores up to 51.0 and vLLM/SGLang serving.
THIS WEEK’S NEWS:
- Pollen Robotics Launches $399 Microduck Biped for Physical AI Training
- OpenAI Uses AI to Design Jalapeño Inference Chip in Nine Months, Beating Nvidia Blackwell
- OpenAI Ends Cursor Contract Following $60 Billion SpaceX Acquisition
- Anthropic Previews Model Hardware Standard for AI Agents to Operate Physical Devices
- Anthropic Merges Claude Chat and Cowork Memory
- Agent Scaffolding Drives Up to 139x Cost Variance in MCP Versus CLI Benchmark
- LAION Releases 10-Million-Hour Open Video Dataset
- Chinese State Hackers Double Cyberattack Volume Using DeepSeek
- OpenAI and 154 Other Organizations Call for Collective Cyber Defense
- Nvidia Agrees to Buy Open Source AI Platform Hugging Face for $12.9 Billion
- NVIDIA AVO Agent Harness Scores 100% on ARC-AGI-3 Public Set
Pollen Robotics Launches $399 Microduck Biped for Physical AI Training
https://pollen-robotics.com/microduck/blog/introducing-microduck/

The News:
- Hugging Face subsidiary Pollen Robotics opened pre-orders for Microduck, a $399 25-centimeter bipedal robot designed as an open-source platform for training and testing physical AI.
- The under-800-gram robot carries 15 motors, an articulated beak for gripping, a small depth sensor, and a camera.
- Unlike the company’s voice-focused Reachy Mini desktop robot, Microduck is built for physical action and can walk, crouch, roller-skate, and get back up from many common falls.
- The software stack is available on GitHub and includes the robot control SDK, a simulator, and reinforcement learning tools for sim-to-real deployment.
- First deliveries for North America, Europe, and the UK are targeted before Christmas.
My take: I was a bit too slow to order this one, and the delivery time for new ducks quickly slipped to 4-6 months. But what you get is a fully assembled walking biped robot for $399 with a complete simulation and reinforcement learning stack already published on GitHub. The Microduck should appeal to anyone teaching, experimenting, or just curious about embodied AI. The main discussion on Hacker News focused mostly on how long the battery lasts and whether Microduck is able to return to the charger on its own. Being able to charge itself would be amazing, and that alone could make this little thing useful for so many use cases! Pollen Robotics hasn’t published any technical details yet, so let’s keep hoping. I can’t wait for the duck to arrive at our office. 🦆
Read more:
- Digital Trends: Microduck Makes Physical AI Training Cheaper
- Hacker News: Discussion of Microduck
- Progressive Robot: Microduck Preorder and Production Caveats
- TechCrunch: Hugging Face Launches the $399 Microduck
OpenAI Uses AI to Design Jalapeño Inference Chip in Nine Months, Beating Nvidia Blackwell
https://openai.com/index/jalapeno-first-results/

The News:
- OpenAI published the first benchmark results for Jalapeño, a custom inference chip it co-developed with Broadcom, reaching tapeout nine months after initial design using its own AI models to accelerate both hardware design and software programming.
- Tested on the public InferenceX benchmark against Nvidia Blackwell systems, the 700-watt chip delivered 1.5 to 1.9 times more throughput per watt and up to 3.6 times lower latency across three large open-weight models.
- For selected GPT-OSS components, AI-generated implementations ran 1.5 to 1.8 times faster than existing human-expert-written implementations, while AI helped the engineering team bring three open-weight models outside the original production plan to high performance within two months.
- Semiconductor research firm SemiAnalysis analyzed the results, noting that Jalapeño uses HBM4 memory and that its single-token prediction throughput per megawatt surpasses Nvidia Rubin’s published results.
- Initial deployment within OpenAI’s compute infrastructure is scheduled to begin by the end of the year, with a second hardware generation already deep in development.
My take: When you read this news, it’s easy to understand why both Anthropic and OpenAI have prioritized software development over writing skills. AI models are now actively building both the hardware and software used to run new AI models, and they are doing it in record time with results far outperforming any similar work done exclusively by humans. The feat itself - building a new chip from scratch in less than one year that outperforms the latest NVIDIA Blackwell chips - is impressive enough. But if anything, this gives us a clear view of where we are heading in the near future. The iterative AI improvement loop has begun, and if we thought progress was going quickly up until now, it will only accelerate going forward.
Read more:
- Futurum Group: Did AI break chip-design timelines?
- Hacker News: Discussion of Jalapeño and AI-designed chips
- SemiAnalysis: OpenAI Jalapeño versus Nvidia Blackwell
- Tom’s Hardware: Jalapeño AI ASIC unpacked
OpenAI Ends Cursor Contract Following $60 Billion SpaceX Acquisition
https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/

The News:
- On August 28, OpenAI announced it intends to wind down its contract supplying AI models to the popular coding platform Cursor, with a proposed shutoff date of November 12, following the tool’s $60 billion acquisition by SpaceX.
- OpenAI stated it used a change-of-control clause to cancel the custom agreement because it lacked confidence that SpaceX would respect its terms of service, citing past contract violations by Elon Musk’s companies.
- OpenAI models currently handle about 5% of user traffic in Cursor, according to co-founder Michael Truell.
- Within hours of the announcement, Anthropic stated it will increase its compute supply to support Claude models inside the editor.
My take: OpenAI is soon to launch GPT-6, and from what I have seen so far, this will be a much bigger leap than we got last year with GPT-5. And it will not be available for Cursor users. While 5% user base might seem like a small number, if you filter out smaller models that require more requests and focus on enterprise users, I think that percentage is much larger. Users can still continue accessing OpenAI models through API access, but good luck using GPT-6 that way when it’s launched next month without spending a fortune. I am convinced this decision will drive enterprises away from Cursor, possibly straight to OpenAI’s own tool, Codex.
Read more:
- CNBC: SpaceX to acquire Cursor for $60 billion
- InfoWorld: Enterprise risks surrounding the Cursor acquisition
- Times of India: OpenAI ends its relationship with Cursor
- X / Fabian Stelzer: Reaction to account-linking friction
- X / John: Reaction to Cursor’s own-model strategy
Anthropic Previews Model Hardware Standard for AI Agents to Operate Physical Devices
https://www.anthropic.com/news/model-hardware-standard-research-preview

The News:
- Anthropic opened a research preview of the Model Hardware Standard (MHS), a shared specification developed with HHMI Janelia Research Campus that lets AI agents safely operate physical lab and manufacturing equipment without weeks of custom integration.
- The standard introduces a standardized driver that uses read and write primitives for devices like microscopes, liquid handlers, and robotic arms.
- In early testing, quantum computing company QuEra used MHS to let an AI agent build a controller that recovered a laser’s precise frequency lock 99.3% of the time without human intervention.
- AWS plans to support MHS through Strands Robots, Doosan Robotics is testing MHS with its robotic arms, and Universal Robots plans to add support to its robotics platform.
- Anthropic is currently inviting stakeholders to join a waitlist for the research preview while it develops physical safety evaluations, with an open-source release planned afterward.
My take: MHS is like MCP but for robots. It shifts the agentic integration work from AI companies to hardware manufacturers, requiring them to expose capabilities through a specific protocol. There are two challenges with getting this protocol into practical use. First, Anthropic clearly writes that “spatial and physical reasoning have limitations that still require expert oversight” in Claude, meaning it lacks spatial intuition. This means that if a robotic arm degrades over time or hits an obstacle, the AI has no idea what just happened. The second challenge is that the entire safety burden for AI agents connected to robots is also shifted over to the hardware manufacturers. This is probably a good thing, but it means they have to consider use cases the robot was not initially designed for. I have no doubt that upcoming models from Anthropic will have full spatial understanding, so give this 6-9 months and I think the interest in this protocol will be substantial from all manufacturers.
Read more:
- AI/TLDR: Model Hardware Standard Overview
- Hacker News: Discussion of the Model Hardware Standard
- The Register: Anthropic Proposes a Hardware Interface for AI Agents
- X / Marcus Kim: Reaction to MHS Device Assumptions
Anthropic Merges Claude Chat and Cowork Memory
https://thenextweb.com/news/anthropic-claude-cowork-shared-memory-default

The News:
- On August 25, Anthropic unified the memory systems for its standard Claude chat and its cloud-based Cowork agent, allowing user context to flow automatically between conversations and task workflows.
- The shared memory is on by default for Free, Pro, and Max users, while Team and Enterprise administrators must explicitly enable it for their organizations.
- Claude now extracts and updates memory topics live while a conversation is still underway, eliminating the previous pause where the system waited until a conversation ended to summarize it.
- Users must manually opt in to save sensitive topics like health, religion, or political beliefs, and the system strictly blocks retention of government IDs and criminal records.
- The update excludes local Cowork sessions and the Claude Code developer tool, with Claude Code remaining on a separate, isolated memory architecture.
My take: Anthropic has three primary tools to interact with their models: Claude Chat, Claude Cowork, and Claude Code. This feature merges the memory between Chat and Cowork, but not Code. In practice, it means that your random brainstorms and quick scribbles now automatically influence how the agents in Claude Cowork format documents and do their research. Some users, like Stefan Galloni on X, are quite vocal and call this a data quality problem, but I’m not too sure. The more Claude knows about you, the easier it can do your job for you. For most users, this will just result in much higher quality work from Claude Cowork. Claude Code is kept outside this merger and still has its own isolated memory, and I hope they keep it that way.
Read more:
- Anthropic Support: Chat search and memory across Claude
- The Register: Claude and Cowork share what they know
- X / Stefano Galloni: Persistent memory as a data-quality problem
- ZDNET: Claude and Cowork share memories unless users opt out
Agent Scaffolding Drives Up to 139x Cost Variance in MCP Versus CLI Benchmark
https://arxiv.org/abs/2608.08654

The News:
- On August 9, researchers published a benchmark comparing AI agent tool use over the Model Context Protocol against standard command-line interfaces, finding that the choice of agent scaffolding drives cost differences of up to 139x while the interface itself has no stable effect.
- Across a fixed Git repository task using five language models and seven agent harnesses, paired MCP-to-CLI cost ratios fluctuated from 0.43x to 29x within the same agents.
- Two of the tested scaffoldings completed the entire task using only standard shell access and were 5.0x to 28x cheaper to run than the five frameworks that support MCP.
- The financial penalty for failure differed by interface, with 12.9 percent of all money spent on MCP runs buying no completed work compared to 2.2 percent for CLI runs.
My take: The important finding from this study is not that CLI can be more efficient than MCP, but rather that if you roll out agents in an enterprise setting, you need to keep strict control of how they fail. Unsuccessful attempts alone consumed over 12% of all MCP-related costs according to this study. The researchers used small models with limited reasoning capabilities, which I am sure drastically affected the difference between MCP and CLI. I don’t think the gap would be anywhere near as large with bigger models with better reasoning. But it again shows that you really need to know what you are doing when rolling out agentic solutions; otherwise, you could end up with a costly system that sometimes performs very poorly without you knowing why.
Read more:
- arXiv: The Scaffold Effect in Coding Agents
- Formal AI: Stop Building Scaffolding
- Hacker News: Discussion of MCP Context Costs
- X / Yunhao Jiao: Reaction to the MCP Versus CLI Benchmark
LAION Releases 10-Million-Hour Open Video Dataset
https://projects.laion.ai/bvd/

The News:
- LAION released LAION-BVD, an open dataset of 80 million videos totaling 10 million hours designed to train multimodal video, audio, and image AI models.
- The collection builds on 1.3 billion URLs collected from CommonCrawl and includes 55 million synthetically captioned video clips alongside 300 million image frames.
- ViCLIP models trained on the dataset match or outperform InternVid-trained models by up to 2.1% on standard video-text benchmarks.
- The dataset is restricted to non-commercial research use and its lightweight versions distribute only URLs and metadata, requiring developers using those versions to download the actual media files themselves.
My take: Using cheap local models, LAION added captions to 55 million videos, which can now be used to train models on the connections between video frames and text. This works in both directions: for AI models learning how to find specific video frames like “an old man repairing a car engine” and for text-to-video models rendering video from text input. The main drawbacks are that you first need to request access to download the 41-terabyte repository, you cannot use it for any commercial purposes, and the caption accuracy in the dataset is not perfect.
Read more:
- arXiv: A Common Pool of Privacy Problems
- arXiv: LAION-BVD Paper
- Hacker News: Discussion of LAION-BVD
- LAION: Dataset Versions
- X / Pedro Anisio Silva: Reaction to LAION-BVD
Chinese State Hackers Double Cyberattack Volume Using DeepSeek

The News:
- On August 24, Taiwanese threat intelligence firm TeamT5 reported that Chinese state-affiliated hacking groups more than doubled their cyberattack volume by using AI, including DeepSeek and other open-source models, to automate reconnaissance and malware development.
- Hackers selected DeepSeek for its low operating costs and open-weight design, which can be run without the provider-side behavioral monitoring used for Western frontier models.
- TeamT5 logged zero incidents using Moonshot’s more capable Kimi K3 model, noting its inference costs are too high to affordably scan hundreds of targets simultaneously.
- In a separate July 30 disclosure, Palo Alto Networks reported that a lone Chinese-speaking operator used DeepSeek inside an open-source agent framework to launch a largely autonomous attack campaign against 460 internet-facing systems.
My take: I think anyone working in AI could see this coming. I am mainly surprised that the attacks only doubled in volume and did not increase more. Maybe this will change with the upcoming models. Current SOTA open-source models like DeepSeek perform well for cyberattacks, but they are not yet on the level of Claude Mythos and the upcoming GPT-6. When they do reach that level, however - which I expect in early 2027 - I think we should expect at least a 100x attack rate increase in a very short timeframe. Which leads us to the next news item.
Read more:
- Aikido: Cyber AI model benchmark
- Cloud Security Alliance: DeepSeek-driven autonomous cyberattacks
- NIST CAISI: Evaluation of DeepSeek AI models
- Palo Alto Networks Unit 42: Chinese-speaking actor’s AI-enabled attack campaign
- X / SumPlus: Reaction to DeepSeek cost efficiency
OpenAI and 154 Other Organizations Call for Collective Cyber Defense
https://openai.com/collective-cyberdefense/

The News:
- OpenAI published an open letter, now signed by 155 organizations, including Microsoft, Google, and Anthropic, calling for a collective global surge in cyber defense before AI-enabled attacks escalate.
- The coalition warns that defenders have a “limited window” to fix longstanding vulnerabilities before increasingly capable AI models make cyber attacks significantly more widespread and sophisticated.
- The letter urges frontier AI companies to provide under-resourced critical infrastructure defenders with responsible model access, funding, and hands-on support.
- It calls on all organizations to make security an immediate leadership priority and asks governments to fund cyber defense for essential services that lack budgets.
- OpenAI’s August 10 expansion of its Daybreak program gives approved defenders access to the specialized GPT-5.6-Cyber model designed to reduce refusals on certain tasks.
My take: I give it 6-9 months before we start to see cyberattacks on a whole new level, largely thanks to next-generation models being able to orchestrate attacks in ways previously not possible. I am just not sure how signing this document with no collective budget or binding commitment will help against that.
The foundation models available to everyone today block almost everything related to cybersecurity, which means they also stop you from properly protecting yourself. We saw that in the Hugging Face incident a few weeks ago. You need specific models like GPT-5.6-Cyber or Claude Mythos 5 to protect yourself, but those models are only available to a few select companies. It’s easy for OpenAI to tell every company to “Make cyber defense an immediate leadership priority” when, in practice, only a select few have access to the tools required to actually make a difference.
Read more:
- CISA: Joint Cyber Defense Collaborative
- Hacker News: Discussion of collective cyber defense
- OpenAI: Expanding Daybreak
- OpenAI: Trusted access for cyber defense
- X / Loopsaaage: Reaction to organizational incentives
Nvidia Agrees to Buy Open Source AI Platform Hugging Face for $12.9 Billion

The News:
- Nvidia has reportedly agreed to acquire open-source AI platform Hugging Face for $12.9 billion, which, if completed, would give the chip maker control over Hugging Face’s model repository.
- The platform is used by more than 13 million developers worldwide.
- Unlike Nvidia’s recent technology licensing deals with AI startups like Groq, a direct purchase requires mandatory US and EU antitrust review.
My take: Hugging Face has over 13 million developers, and NVIDIA now gets a direct connection with all of them. I fully expect NVIDIA to continue improving Hugging Face as a platform, but I also expect them to invest further in making CUDA workloads the easiest and best-performing alternative. With new chips like OpenAI Jalapeño being developed in months instead of years, cheaper and better-performing alternatives to NVIDIA chips could arrive en masse in the next 3-5 years. But if NVIDIA owns the distribution platforms, they can also control which models they provide.
Read more:
- CNBC: Nvidia Reportedly Agrees to Buy Hugging Face
- Hacker News: Discussion of Nvidia and Hugging Face
- LinkedIn / John Cloud: Nvidia Is Buying Distribution
- Nvidia Newsroom: Nvidia and Hugging Face DGX Cloud Partnership
- TechCrunch: Nvidia Closes In on Hugging Face Acquisition
NVIDIA AVO Agent Harness Scores 100% on ARC-AGI-3 Public Set

The News:
- On August 21, NVIDIA reported that its Agentic Variation Operators (AVO) system scored a perfect 100.00 on the ARC-AGI-3 public set by wrapping Claude Opus 5 in a long-horizon memory and supervision harness.
- Claude Opus 5 scores about 30% on ARC-AGI-3 in a separate evaluation that is not directly comparable with the AVO run.
- The system completed all 183 public levels using 6,624 environment actions, which is 12% fewer than VISTA.
- NVIDIA originally built AVO for autonomous software engineering, where it ran for seven days to produce multihead attention GPU kernels up to 10.5% faster than FlashAttention-4.
- According to AIHOT, François Chollet praised the work but noted the 100% score covers only the public demonstration set, which he compared to clearing a video game’s tutorial level.
My take: ARC-AGI-2 made sense to me since it measured how well a model could reason to solve problems. ARC-AGI-3 seems flawed in that it does not include the harness in the benchmark. If you have been following the evolution of AI models over the past few years, you know it went something like this: Test Time Compute -> Reinforcement Learning with Verifiable Rewards (RLVR) -> Harness optimizations.
Today, how a model performs depends heavily on how the agentic harness is configured, making tests like ARC-AGI-3 - which use a “generalized harness” by default - such a poor reference. By modifying just two settings in the harness, OpenAI managed to triple their score. And here, NVIDIA scored 100% using its own agentic harness. If we are going to use benchmarks like ARC-AGI-3 to see how well models can play games, I feel it’s only fair that model providers can use their own harnesses.
Read more: