Tech Insights 2026 Week 40
September 28, 2026
Last Tuesday, Nat McAleese at Anthropic wrote on X: “Opus 5.5 is way, way, way better than Opus 5. Sorry about that model, please try this one”.
To put that quote in perspective, Anthropic has really struggled over the past few years with their models. Opus 4.0 was a great model in May and June 2025, but once interest outgrew supply, they had to scale down the harness and model to make it accessible to more people. Opus 4.6, released in February, managed to get some of the quality of the initial 4.0 release back through optimizations and new infrastructure. But demand rose exponentially again, and Opus became worse with every release, ending up with Opus 5, which was a behavioral disaster.
Opus 5.5 has impressed me in ways that very few models do these days. It’s fast, it’s relatively cheap, and it mostly performs better than Fable and Astra on programming, creative tasks, writing, and visual design. On the GDPval-AA v2.1 professional work evaluation, Opus 5.5 scored 1846 Elo, ahead of Fable 5.1 at 1735 and GPT-6 Astra at 1542, at a base price of $4/$20 per million input/output tokens. Where I would burn through a full weekly Max-20 account with Astra in less than 24 hours, Opus 5.5 seems to last the entire week despite having it work around the clock on agentic tasks.
To get a sense of what Opus 5.5 is capable of, take one minute of your day and checkout the video below. Claude Opus 5.5 drew all 7,200 frames and composed the entire soundtrack. No downloaded footage, no game clips, no music files. It was posted by @Prasenjit on X on September 26.
Full prompt to generate this video: “Make a cinematic 2-minute video on the history of video games. Make it epic and emotional, like a documentary trailer. Go all out. Compose the soundtrack yourself, fully synthesized in code. No samples, no audio files, no VST instruments. Cut the visuals to the music. 1920x1080, 60fps. No downloaded images or footage, every frame generated in code. Render the final video with the music mixed in as an MP4 and save it to Desktop.”
A few other examples worth checking out (on X):
- A music video about the life of CodeRabbit’s rabbit, made with Opus 5.5 in one shot.
- Motion graphics by Mirage, made in one shot with Opus 5.5 and Tesseract.
- A video by Zsolt Kacso, made with just two prompts with Opus 5.5.
Six months ago, most people seemed to agree that it would take a very long time before an AI model could replace the roles of art directors, motion designers, or software architects. Now, with Opus 5.5 producing videos like the ones above and improving the performance of claude.ai over 3 times by itself, this has all changed. Models like GPT-6 Astra and Fable 5.1 gave you an insight into how a future could look when AI models perform really advanced tasks at a significant price. Opus 5.5 gives you that future today at a reasonable price.
Thank you for being a Tech Insights subscriber!
Listen to Tech Insights on Spotify: Tech Insights 2026 Week 40 on Spotify
Notable model releases last week:
- CLM-8B by Contrastive-LM. Fast decision model that helps AI agents pick the best next action, tool, or candidate answer from a list. Built on Qwen3-8B, it is free to download and runs on one gaming GPU.
- FLUX 3 Action by Black Forest Labs. Model that controls robot arms from camera images and instructions, built on FLUX 3’s video training and also aimed at simulators and games. The 7B weights are free to download and lead the RoboLab robot benchmark.
- Gemini 3.8 Flash TTS and Flash-Lite TTS by Google. Updated text-to-speech models that can now design new voices from a text prompt or clone one from a 30-second sample, for audiobooks, podcasts, dubbing and voice agents. Flash TTS ranks first on Hume AI’s Voice Design Benchmark.
- Grok 4.7 by SpaceXAI. Update built on a larger base model that handles long coding and knowledge-work tasks better and checks its own work more carefully. It keeps Grok 4.6’s price and is available in Cursor and the Grok API.
- Grok Voice Transcribe 2.0 by SpaceXAI. Speech-to-text update for transcribing support calls, videos, and voice commands, now twice as accurate as version 1.0 at the same price. It ranks first for accuracy among streaming models on the Artificial Analysis leaderboard.
- Hemmingway-1 by Altworld. Writing model that drafts everyday messages, emails and awkward notes in a natural human voice, without extra commentary. Built on Qwen3.8-27B, it is free to download for non-commercial use.
- MiMo-V2.6 by Xiaomi. Free-to-download models that take text, images, audio and video, mainly for coding and agent tasks. The Pro version is the top open model on the Artificial Analysis Intelligence Index, at unchanged API prices.
- SAM 3.1 by Meta. Model that finds, outlines and tracks objects in images and video from a short text prompt, with no training data needed. It is now a hosted service on Meta Model API, at $2.50 per 1,000 images.
THIS WEEK’S NEWS:
- Anthropic Launches Claude Opus 5.5 With 40% Cheaper Workloads
- Anthropic Expands Claude Marketplace With 2,000 Connectors and Partner Software
- Anthropic Merges 3,000 AI-Written Changes to Make Claude.ai 3x Faster
- Meta Plans Muse Agent for Smart Glasses and Announces $1,299.99 VR Headset
- Meta and Google Introduce Real-Time Video Avatars for Voice Agents
- Google Launches Project Suncatcher Prototype to Test Orbital AI Data Centers
- OpenAI Releases GPT-6 Sol and Luna With 50% Lower API Prices
- OpenAI Releases MentalHealthBench to Evaluate AI Responses in 1,215 Conversations
Anthropic Launches Claude Opus 5.5 With 40% Cheaper Workloads
https://www.anthropic.com/claude-opus-5-5

The News:
- On September 22, Anthropic released Claude Opus 5.5, which the company says delivers Claude Fable 5.1 capabilities at a 40% lower typical compute cost than Opus 5.
- Base API pricing drops to $4 per million input tokens and $20 per million output tokens, alongside a steep cache read reduction to $0.20 per million tokens.
- On the GDPval-AA v2.1 professional work evaluation, Opus 5.5 achieved an 1846 Elo score, outperforming Fable 5.1 at 1735 and OpenAI’s GPT-6 Astra at 1542.
- Because its unmitigated cybersecurity capabilities are comparable to Claude Mythos 5.1, Anthropic forces most cyber-related prompts to transparently fall back to the older Claude Opus 4.8.
- Endor Labs tested the model and found it finishes coding tasks up to twice as fast as Fable 5.1, though researchers noted its security pass rate dropped 19 points when they removed tasks the model had memorized from its training data.
My take: When using a model like Claude Opus for programming, most of the cost comes from cached tokens due to the way the model iterates on the conversation before producing output. And here comes Opus 5.5, where the cost of cached tokens was cut by 60%. I have been running Opus 5.5 non-stop, 24 hours per day since its release, and I am just approaching the limits of my Max20 subscription. The performance utilization of this model is incredibly good. Several users have posted cost comparisons on X, such as @HarshithLucky3, who demonstrated that a 3D model of a Waymo using 41.2 million cached tokens cost just $27.39 to build with Opus 5.5, compared to $50.56 with Opus 5.
I have been comparing Opus 5.5 with GPT-6-Astra and Fable 5.1 over the past week, and so far, I like Opus 5.5 more than the other two models. Not only is it significantly cheaper and faster, but I also find it can reason better, has better creativity, and can write better text. It is just a much better model. I did not expect this, and I have not used an Opus model since version 4.6 in February. But Opus 5.5 is different. It is without a doubt the best model I have ever used, surpassing both Astra and Fable when it comes to creativity and quality of results. And it will probably stay that way until at least Tuesday, when OpenAI holds its DevDay for 2026.
Read more:
- Endor Labs: Opus 5.5 is cheaper and faster, but memorization keeps it off the top spot
- Hacker News: Discussion of Claude Opus 5.5
- SiliconANGLE: Anthropic releases Opus 5.5 and OpenAI counters with cheaper GPT-6 models
- X / Token Gremlin: Reaction to Opus 5.5 cache pricing
- Zvi Mowshowitz: Claude Opus 5.5 system card review
Anthropic Expands Claude Marketplace With 2,000 Connectors and Partner Software
https://claude.com/blog/claude-marketplace

The News:
- On September 23, Anthropic expanded the Claude Marketplace to consolidate more than 2,000 plugins and connectors alongside third-party AI agents, software products, and consulting services into a single enterprise platform.
- Enterprise teams can purchase AI software from partners like CrowdStrike, Cursor, Harvey, and Snowflake using a portion of their existing Anthropic spending commitments.
- Early customers including CodeRabbit and ThoughtSpot have redirected portions of their Anthropic commitments to fund Vercel and Snowflake infrastructure through the platform.
- The directory lists consulting firms from the Claude Partner Network, such as Accenture, Boston Consulting Group, and Deloitte, to handle enterprise deployments.
- Adding a custom connector does not trigger a new FedRAMP review for public sector users, as the execution infrastructure is covered by Claude for Government’s existing FedRAMP High authorization.
My take: Let’s say you are a manager at a large company, and many of your employees use Claude. Now you want to buy a Claude-powered tool. Typically, it would involve a new contract, a new security review, and a new budget proposal, which could take months. Now, with the Claude Marketplace, enterprise customers can use a portion of their committed Anthropic spend to buy Claude-powered products straight from companies like CrowdStrike, Cursor, Harvey, Legora, and Snowflake. The security review may still apply, but the contract and budget are already in place.
If you are creating plugins or connectors for Claude, then the Claude Marketplace is the place to be. Anthropic still has a waiting list to get approved, but once that loosens up I think this will grow to tens of thousands of plugins and connectors in just a few months. The Claude Marketplace also shows how Anthropic, OpenAI, Google, and Microsoft are all building their own unique value offerings on top of their models. Ease of access is going to be key to success, and I bet OpenAI will announce something similar very soon, maybe as early as DevDay on Tuesday.
Read more:
- Anthropic: MCP 2026-07-28 spec brings a stateless core to Claude
- Azalio: Claude Marketplace targets AI procurement bottlenecks
- Harmonic Security: Securing Claude Cowork, a security practitioner’s guide
- Remio: Claude Marketplace turns AI distribution into the next enterprise contest
- X / Jay Song: Reaction to Plugin4Shell and plugin SHA pinning
Anthropic Merges 3,000 AI-Written Changes to Make Claude.ai 3x Faster
https://claude.dev/blog/how-we-made-claude-ai-faster/

The News:
- On September 23, Anthropic reported using an internal research model roughly comparable to Opus 5.5 to merge more than 3,000 code changes in a two-week sprint that made claude.ai about three times faster.
- At the 75th percentile, the time to a typeable page on a fresh load of claude.ai dropped from 3.1 seconds to 0.55 seconds.
- Engineers managed the project from a single Slack channel, where Claude analyzed telemetry, built deterministic benchmarks, and opened pull requests for human review.
- The model optimized JavaScript instruction counts and rendering paths, finding and fixing bottlenecks like a syntax-highlighting regex that stalled on non-Latin characters.
- Anthropic says it shipped the optimizations without a single customer-facing incident or rollback by using automated continuous integration guardrails and nearly 200 short-lived feature flags.
My take: This is the way you approach performance optimizations today. Pick an AI model and let it measure things like React commits and style recalculations. Then set a goal and let it try every different approach that improves these values. I did the same with my plugin Notebook Navigator: I asked Claude Fable 5 to measure everything that affected startup times and rendering performance, quickly discard paths that were irrelevant, then try every possible variation to improve things that actually made a difference, all while keeping a logbook on how things improved. The end results in my case were 70% shorter startup time and twice the rendering performance.
In this specific case, Anthropic says they merged 3,000 changes to claude.ai without a single customer-facing incident. If you want to learn more how this was possible, @lauren (Grok developer at SpaceXAI) last week posted a video on X where she goes into details how she alone was able to ship over 2,500 pull requests in just one month straight to production. We are living in wild times.
Read more:
- Fortune: AI coding agent horror stories and the Amazon outages
- Moderne: AI didn’t break coding, it broke code review
- RuntimeWire: Anthropic says Claude made its apps 3.1x faster across 13 measurements
- Superpower Daily: What the 3.1x figure does and does not show
- The Pragmatic Engineer: Are AI agents actually slowing us down?
Meta Plans Muse Agent for Smart Glasses and Announces $1,299.99 VR Headset
https://www.meta.com/en-gb/blog/meta-connect-2026-everything-we-announced/

The News:
- On September 23, Meta announced that its Muse personal AI agent is coming to its smart glasses, added computer use to Muse for Mac, and introduced a 100-gram spatial computing headset arriving next spring for $1,299.99.
- Muse can now drive native Mac applications in the background, utilize a real-time voice mode, and access new commerce connectors for retailers like Walmart and Best Buy.
- The company’s smart glasses will allow wearers to activate the agent by name, letting Muse identify and act on physical objects in the user’s field of view without requiring verbal descriptions.
- A software update launching later this year will add an FDA-cleared hearing enhancement feature to supported AI glasses for a one-time $149.99 fee.
- The flagship Meta VR Glasses feature a 5K micro-OLED display and move the battery and processing hardware into an external puck to reduce the headset weight to a fifth of the Meta Quest 3.
My take: Muse is quickly becoming Meta’s version of Microsoft Copilot. It’s everything and everywhere. Muse will be available in your AI glasses, on your Mac laptop, on your phone, and in your portable, always-connected Muse charm device. The new strategy for Meta seems to be that everyone should send everything they see and hear to Meta all the time and everywhere.
Meta’s core business requires them to know as much about you as possible so they can sell ads to companies with the best possible hit rate. Reviewers cited by Dataconomy have already reported that Muse reads private notifications without explicit permission and keeps pushing users to connect their email and banking data. Maybe it’s because I am from the EU that I see more problems than possibilities with Meta and Muse. Either way, it will be interesting to follow adoption going forward.
Read more:
- ABC News: What to know about Meta’s Muse AI agent
- Dataconomy: Muse reportedly read private notifications without permission
- Forbes: Amazon blocks Meta’s Muse agent from shopping on Amazon.com
- Hacker News: Discussion of Muse’s 6.8GB filesystem export
- Meta for Developers: Meta Connect 2026 recap
Meta and Google Introduce Real-Time Video Avatars for Voice Agents
https://research.meta.ai/blog/bringing-your-muse-to-life https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-with-live-avatar/

The News:
- On September 23 and September 24, Meta and Google introduced live video avatars that synchronize with their real-time voice models, with Meta previewing Muse Realtime Avatar for its consumer agent and Google releasing Gemini 3.8 Live Avatar for enterprise customers.
- Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise, where it can execute background tool calls and transition across 97 languages without visual drift.
- Meta’s Muse Realtime Avatar streams 448x768 video at 25 frames per second with approximately 870 ms of latency, supporting 12 concurrent sessions on a single GB200.
- To reach interactive speeds, Meta distilled a 40-step diffusion teacher (120 model evaluations per chunk) into an unguided two-step student model.
- Both companies embed invisible watermarks into their generated video streams, using SynthID for Gemini and Video Seal for Muse.
- In live evaluations run by Meta, human raters preferred Muse Realtime Avatar over commercial competitors Runway Characters and HeyGen LiveAvatar.
My take: I think these virtual avatars look truly horrible, and I cannot for the life of me understand why companies keep pushing them into their services. Every month that passes where Anthropic and OpenAI do not add visual avatars to their products is a good month. In their current state, digital visual avatars only distract from the discussion and give viewers an odd, uncomfortable feeling, especially when they try to mimic actual people. It’s telling that Meta’s own evaluation only compared its avatar against other avatars (Runway and HeyGen) and never against voice-only interaction, I actually think most people would prefer the voice-only option here.
Read more:
- Latent Space: Meta Connect 2026 roundup on Muse voice, video and glasses
- Orca Router: Muse Realtime Avatar and the 870 ms question
- Remio: Gemini 3.8 Live Avatar is GA, and the real test is production trust
- The Register: Google Gemini can present cartoon or lifelike avatars
- X / Trisha Koshy: Reaction to avatars raising the bar for disclosure
Google Launches Project Suncatcher Prototype to Test Orbital AI Data Centers

The News:
- Next week, Google will launch a prototype satellite built with Planet Labs to test whether its Tensor Processing Units can survive the radiation and thermal extremes of low Earth orbit as part of a proposed space-based AI data center.
- The long-term architectural design envisions clusters of 81 satellites flying in close formation separated by 100 to 200 meters, communicating via free-space optical lasers.
- Ground tests ahead of the SpaceX Transporter-18 rideshare launch showed that Trillium TPUs survived a radiation total ionizing dose equivalent to a five-year mission without permanent failure.
- An accompanying Google preprint paper calculates that if launch costs drop below $200 per kilogram by the mid-2030s, the cost of launching and operating a space-based data center could become roughly comparable to a terrestrial data center’s energy costs per kilowatt-year.
- Google plans to launch two additional satellites in early 2027 to test the laser interconnects.
My take: Orbital data centers only start to make sense if the cost to put a kilogram of material into space drops below $200. The Space Shuttle cost tens of thousands of dollars per kilogram to put into orbit, and the reusable Falcon 9 rocket is down to $3,600. Now, even at $200, Google estimates that putting a data center into space would just match the energy bill of a data center on the ground, not the cost of building and running one. We have a long way to go before it’s profitable to put data centers in space, if we ever get there.
Read more:
- Google Research: Exploring a space-based, scalable AI infrastructure system design
- Hacker News: Discussion of Google’s Project Suncatcher
- LinkedIn / Jonathan Trent: Post about solar and cooling panel area for orbital data centers
- TechCrunch: Why the economics of orbital AI are so brutal
- The Verge: Billionaires want data centers everywhere, including space
OpenAI Releases GPT-6 Sol and Luna With 50% Lower API Prices
https://openai.com/index/introducing-gpt-6-sol-and-luna/

The News:
- On September 22, OpenAI expanded its latest model family with GPT-6 Sol and GPT-6 Luna, two new mid-tier and budget models that cut API prices by 50% compared to their predecessors.
- GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, and GPT-6 Luna costs $0.10 for input and $0.50 for output.
- Both models support a 90% discount on cached input reads, dropping Sol’s cached input rate to $0.20 per million tokens and Luna’s to $0.01.
- On AutomationBench, OpenAI reports that GPT-6 Sol at xhigh effort scores 33.2% at $0.27 per task, beating Claude Opus 5 at max effort which scored 26.9%.
- The models are available now in ChatGPT Work, Codex, and the API, but neither is accessible in the standard Chat interface.
My take: GPT-6-Sol is the worst model release from OpenAI in a very long time, mainly because most people were expecting it to be an improvement over GPT-5.6-Sol. But it’s not; it’s a degradation. OpenAI won’t confirm this, but I am close to 100% certain that GPT-6-Sol is just an improved version of GPT-5.6-Terra, and not an improved version of GPT-5.6-Sol. So why do I think this?
First - and you can try this yourself - open the ChatGPT app on your phone, switch to “Work” mode, and change the model to GPT-6-Sol. Use “balanced mode” or “medium”; the results are the same. Ask it, “How many weekdays have the letter d in them?” and you will get, “Three: Monday, Wednesday, and Friday.” I tried and verified it myself, and many X users like Gael Breton have posted about it as well. If you have a note in your memory about the number of weekdays or have asked the question before, you will get the right answer. GPT-5.6-Sol does not answer like this, but GPT-5.6-Terra was one of the few modern models that did.
Secondly, GPT-5.6-Terra is now removed and no longer available. And just by coincidence, the pricing for input tokens for the “new and cheaper” GPT-6-Sol is the same as GPT-5.6-Terra. OpenAI has already posted on X that users should expect to see new models at DevDay on Tuesday, and my best guess is that they will launch GPT-6.1-Astra, which will actually be GPT-5.6-Sol distilled from the larger GPT-6-Astra. This should mean it will be as cheap as Opus 5.5 and probably outperform the big, original GPT-6-Astra. I’m a huge fan of the GPT models, and I understand the reason to name models based on their performance, but in this case, they should have learned from Opus 5 and just kept the name Terra.
Read more:
- DataCamp: GPT-6 Sol vs Claude Opus 5.5
- Hacker News: Discussion of GPT-6 Sol and Luna
- Kingy AI: GPT-6 Sol and Luna specs, benchmarks and pricing
- OpenAI: API pricing
- X / NovaXCode: Reaction to GPT-6 Sol as a downgrade
OpenAI Releases MentalHealthBench to Evaluate AI Responses in 1,215 Conversations
https://openai.com/index/introducing-mentalhealthbench/

The News:
- On September 23, OpenAI released MentalHealthBench, an open benchmark of 1,215 synthetic mental health conversations paired with grading rubrics co-created with more than 80 licensed clinicians.
- The dataset covers everyday, high-acuity, and emergency scenarios, using GPT-5.6 Sol to grade responses against expert criteria weighted from -10 to +10.
- OpenAI’s results show steady improvement across model generations, with more advanced models better at seeking context appropriately.
- The benchmark measures coverage of the grading rubric rather than proven clinical benefit.
My take: Benchmarks tend to have the side effect that model vendors start optimizing for them. And in this case, it’s clear to see what will be rewarded. GPT-6 Astra scored 57.3% on MentalHealthBench, where clinicians’ own reference answers scored 38.5%.
Real psychologists write short and focused replies, whereas a language model tends to write extensive answers that check most of the test boxes. While useful for tracking a model’s knowledge, this isn’t the same thing as being helpful to another person. The main risk here is that now that the benchmark is out in the open, most labs will start optimizing for it by just responding with longer and more clarifying answers. For people under stress, even if shorter answers cover less, they might still be better than the most extensive answers.
Read more: