Tech Insights 2026 Week 37

September 7, 2026

Tech Insights 2026 Week 37

The world changed last week.

I spend dozens of hours every week working with the latest AI models, reading about how companies use them, and researching where they are heading. Yet I found myself completely unprepared for the release of GPT-6 Astra last week. I knew OpenAI was going to release it, and I knew it would be good at computer use. But I had no idea just how good it was going to be.

The thing is, it’s almost too good. If I could borrow just two minutes of your time today, please stop what you are doing and visit the links below. Then get back here and continue reading. I promise you haven’t seen anything like this before.

Absolutely mind-blowing examples: Apple Park - built 100% by prompting GPT-6 Astra Matt Shumer on X: Voices from the living room

Game examples: Anshu on X: Voxel game mihawkxxx on X: Roblox fishing game

3D production examples: Sharif Shameem on X: Palace of Fine Arts

2D illustration example: bluedev on X: Xbox Series X controller

With GPT-6, codenamed “Astra”, AI models have now evolved from “large language models” (LLMs) into something else. The size increase and the extensive training in computer use have made GPT-6 Astra transition from something that is good at writing text and source code, into something that can reliably work with any application on your computer with expert skills in virtually any domain.

I’ve had GPT-6 running around the clock since its release, and my personal reaction is something close to shock. This model can literally do everything on your computer that the best humans in any niche can do. Suddenly you have the world’s best Excel user, the world’s best 3D modeler, the world’s best illustrator, the world’s best software architect, and the world’s best project manager, all at your disposal through the Codex app.

Looking forward, I predict a major shift across the entire IT industry, much like what is happening right now in programming. It won’t come overnight, and for some companies it will take years before they even begin. But I am sure it will come. The shift will be a transition from being good at working in different computer programs like Altium Designer, AutoCAD, SolidWorks, Maya, Power Automate, Illustrator, Jira, Microsoft Project, Terraform, Kubernetes, MATLAB, and Dynamics, to becoming good at telling an AI what to do. Your experience in knowing how to do things will mean nothing in this future.

The thousands of hours you have spent learning every shortcut and sequence to get work done quickly in a specific tool will have no purpose when an AI agent does the work 10 times faster. It will be challenging times for everyone, including software developers, when clients can suddenly shift between Blender, 3DSMax, Maya, Houdini, and Cinema 4D depending on current needs, without having to re-learn every workflow specific to each program.

In a not-so-distant future, the difference between a good consultant and a mediocre consultant will be who is best at prompting an AI to get the best possible results at the lowest possible token cost. I predict this will be the same for every single domain working with desktop computer software today. If you spend most of your days comfortably working in a desktop computer application, I strongly recommend that you sign up for an OpenAI subscription and try to become efficient at controlling an AI to do your job. Because this is what most of your days will probably look like in just 1-2 years.

Thank you for being a Tech Insights subscriber!

Listen to Tech Insights on Spotify: Tech Insights 2026 Week 37 on Spotify

Notable model releases last week:

  • Claude Fable 5.1 and Claude Mythos 5.1 by Anthropic. Identical models with general and trusted-access safeguard tiers, scoring 55.8% and 60.9% on Terminal-Bench 4.0 versus Fable 5’s 42.0%, with cache-read pricing cut 75%.
  • Flow Reasoning Models by Alec Helbling et al. Recurrent flow-based structured reasoners using Fixed-Point Forcing for iterative parallel refinement, matching 98.7% on Sudoku-Extreme with 44x fewer inference FLOPs than the next-best method.
  • Gemini 3.8 Flash by Google DeepMind. Outperforms 3.7 Flash on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark, scores 54.9% on HLE-Verified, and supports 1M-token input, 64K output, and computer use.
  • MAI-Transcribe-2 by Microsoft AI. Speech recognition model with diarization and word-level timestamps, averaging 5.2% WER across 60 FLEURS languages, processing audio up to 10x faster than GPT-Transcribe, at $0.10 per hour.
  • Muse Spark 1.3 by Meta. Agent model for long-horizon coding and multi-workflow threads, using ~20% fewer tool calls and ~25% fewer tokens than 1.2, available through Meta Model API.
  • Muse Voice Transcribe by Meta. Real-time streaming audio perception model combining ASR, endpointing, and diarization for 20+ speakers, ranking first on Artificial Analysis streaming speech-to-text with 3.1% WER.
  • Qwen3.8-Max-0902 by QwenCloud. Multimodal agent model with 1M context and 131K max output, adding stronger engineering-scale coding, multi-tool orchestration, and native vision understanding.
  • TimesFM-3 by Google. 330M decoder-only time-series foundation model for zero-shot multivariate forecasting, adding single-pass decoding and ranking first among pre-trained models on Gift-Eval, FEV-Bench, and Time.

THIS WEEK’S NEWS:

  1. OpenAI Releases GPT-6 Astra
  2. OpenAI’s GPT-6 Astra Scores 99.9% on ARC-AGI-3 Benchmark
  3. OpenAI’s ChatGPT Ads Reaches $1 Billion Annualized Revenue Run Rate in Under 200 Days
  4. Anthropic Releases Claude Commerce Agents Blueprint
  5. Cursor Releases Self-Hosted Machines for Internal Cloud Agent Execution
  6. World Labs Announces Atlas to Generate and Reconstruct 3D Worlds
  7. Runway Solaris Generates Real-Time App Interfaces Frame by Frame Without Code
  8. OpenClaw 2.0 Launches With 16,000 Pull Requests and Multiplayer Cloud Sessions
  9. European Commission Designates ChatGPT a Very Large Search Engine Under the DSA

OpenAI Releases GPT-6 Astra

https://openai.com/index/gpt-6-astra/

The News:

  • On September 3, OpenAI released GPT-6 Astra, a multi-step model built for autonomous computer use and coding that scored 72.6% on the offline subset of the OSWorld 2.0 desktop benchmark, using partial scoring.
  • In OSWorld 2.0 latency simulations, the model completes computer-use tasks in about 47% less time than its predecessor GPT-5.6 Sol, at roughly 40 minutes per task.
  • Astra is the company’s first model to reach the “Critical” cybersecurity capability threshold, having successfully discovered two previously unknown zero-day vulnerabilities during internal testing.
  • Standard API access costs $10 per million input tokens and $50 per million output tokens. The model has a 1.05-million-token context window.
  • While OpenAI claimed a 99.9% score on the ARC-AGI-3 benchmark using a custom test harness, the independent ARC Prize organization measured the model at 62.7% on its standard evaluation.

My take: Astra is OpenAI’s first model to be classified as Critical in their cybersecurity rating: “GPT-6 Astra can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step.” In practice, this means that when you ask the model to perform a task, it will quite often think “outside the box” and try different approaches to solve it. I have been running GPT-6 around the clock since release, and this creativity, if you can call it that, is emergent in everything GPT-6 Astra does. Ask it to implement a UI feature, and it will come up with dozens of different approaches to stress-test it that would be difficult to cover as a human.

To me, GPT-6 represents the third major leap in generative AI since GPT-3 was launched in 2020, and Opus 4 and GPT-5 were launched in 2025. LLMs have evolved from simple chatbots with no internet access and no reasoning, to agentic systems with tools, internet access, and reasoning, to GPT-6, which can do any task in any application on your computer with expert skill. This is exponential development at its best, and even if we know it’s happening, it’s still so hard to grasp when it actually is happening.

While agentic harnesses and reasoning models are starting to cause quite a stir in the developer community, I believe it will take 6-9 months before most companies even understand what’s possible with GPT-6. One office worker can now do the job of 20 people, with skills as varied as systems architects, CAD engineers, 3D modelers, shader programmers, Excel experts, Power BI experts, and much more. You get all these skills at your disposal for any task you can imagine, right on your computer. And this is what will start changing the world.

Read more:


OpenAI’s GPT-6 Astra Scores 99.9% on ARC-AGI-3 Benchmark

https://arcprize.org/blog/astra

The News:

  • On September 3, OpenAI released GPT-6 Astra, a new flagship model that achieved a 99.9% score on the Semi-Private set of the ARC-AGI-3 reasoning benchmark when using a provider-specific adapter.
  • Under ARC Prize’s standard provider-neutral testing harness, the model scored 62.7% on ARC-AGI-3 Semi-Private.
  • In the Provider Adapter harness at max reasoning, Astra surpassed the human baseline for action efficiency, using fewer actions than the median human tester on 96% of levels.
  • Despite the near-perfect adapter score, the ARC Prize creators emphasized that saturating this benchmark does not prove Astra has achieved AGI.
  • According to Codersera, base API pricing is $10 per million input tokens and $50 per million output tokens, with a 1,050,000-token context window.

My take: Six months ago, ARC-AGI-3, the latest “IQ test” for AI models, was launched with ARC Prize saying that “Humans score 100%. Frontier AI scores 0.51%.” Well, now both humans and AI models score 100%. The test itself, it turns out, was constructed in a way that was counter-intuitive for today’s models and actually penalized them by not allowing them to remember anything from previous play sessions, meaning AI models did not have the same benefits as humans. Some of the 100 levels were almost impossible to solve if you did not understand the logic from previous levels. Despite this, with the limited brain-washing harness, GPT-6 Astra scored a massive 62.7%. But with the harness allowing the model to remember what it had played previously, like humans, it saturated the test. It’s hard to overstate just how big a leap GPT-6 Astra is; it’s one of the most important AI releases in modern history.

Read more:


OpenAI’s ChatGPT Ads Reaches $1 Billion Annualized Revenue Run Rate in Under 200 Days

https://openai.com/index/expanding-access-to-ai-with-chatgpt-ads/

The News:

  • On August 31, OpenAI announced that its ChatGPT Ads platform reached a $1 billion annualized revenue run rate in less than 200 days and expanded self-service purchasing to Europe, India, the Middle East, and North Africa.
  • The advertising system shows labeled sponsored cards to logged-in users on the Free and lower-tier Go plans, keeping Plus, Pro, and Enterprise accounts ad-free.
  • OpenAI shifted the platform from cost-per-thousand impressions to cost-per-click pricing, with bids reportedly between $3 and $5.
  • The ads are placed below the chatbot’s responses and use the context of the current conversation to target users without giving advertisers direct access to chat logs.
  • Despite the rapid revenue growth, some advertising agencies report that the platform’s dashboard shows clicks that their own analytics tools cannot find.

My take: For countries that cannot afford to pay for AI, ad-based services are absolutely the way to go. OpenAI only shows ads as sponsored cards to logged-in users on the Free and lower-tier Go plans, leaving Plus, Pro, and Enterprise accounts ad-free. Advertisers do not receive any chat logs or personal data from users; instead, OpenAI uses conversation context and user activity to select which sponsored cards to show. OpenAI also says advertising does not influence ChatGPT’s answers.

Read more:


Anthropic Releases Claude Commerce Agents Blueprint

https://claude.com/blog/the-anatomy-of-effective-commerce-agents

The News:

  • On September 2, Anthropic released a blueprint for building Claude-powered shopping and merchant agents, including reference code for retail, travel, telecom, and ticketing systems.
  • The architecture uses a single agent loop equipped with modular skills rather than multiple subagents, which Anthropic says preserves conversation history while reducing latency and cost.
  • The agent cannot execute payments or change live listings directly; the model only proposes changes and emits UI schemas, while the host application enforces approvals and business limits.
  • Anthropic reports retailers running shopping agents on Claude saw carts up to 35% larger and shoppers 60% more likely to complete a purchase.
  • The system is designed to achieve 90 to 99% prompt cache hit rates, with cached input tokens priced 90% below fresh input for standard models and 97.5% below for Claude Fable 5.1.

My take: A commerce agent is “an agent that simplifies buying and selling across an online catalog. Some agents face consumers: they search, compare, substitute, and assemble the order. That could be a retail cart, a travel itinerary, a mobile plan change, or seats held for a show. Some agents face the business: they answer questions about sales, run promotions and campaigns, and manage inventory and pricing.”

If this is something you are doing, you should really read through this document. It covers different strategies like system prompts vs. skills, agentic tooling, latency, prompt caching, model choice, memory, and much more. If you want to own your data and customer experience, this is probably something you will want to build. Just be aware that if your team has not worked with AI agents on this level before, it’s quite the learning journey.

Read more:


Cursor Releases Self-Hosted Machines for Internal Cloud Agent Execution

https://cursor.com/blog/self-hosted-machines

The News:

  • Cursor launched Self-Hosted Machines, a deployment feature that lets enterprise teams keep their codebases and tool execution on their own infrastructure while Cursor handles the agent inference in the cloud.
  • The architecture moves only the execution environment to a user’s network or a sandbox provider like AWS Lambda or Vercel.
  • Enterprise administrators can configure shared worker pools that automatically scale with developer demand and hibernate idle machines to save compute costs.
  • The self-hosted Linux and Mac workers support computer use, allowing agents to control local browsers and take screenshots inside company networks.
  • Team pools require a Cursor Enterprise plan, while individual developers can connect a single personal machine.

My take: If you have ever tried to employ a somewhat complex agentic setup with a third-party agent provider, you know how tricky it is. If you run agents that are hosted at the provider, you basically need to open up your entire internal infrastructure to another company. And then you encounter issues when the AI agents need to access things like internal source control repos or internal web browsers without letting the agents in through a VPN.

Cursor Cloud Agents are used to delegate coding tasks and get back a reviewable PR. There are dozens of alternatives to this, but if you use the Cursor environment, you are most probably already using this as part of your infrastructure. This update simply makes it possible to run the agents inside your own network, which gives a safer and more flexible environment. The workers connect to Cursor over outbound HTTPS, so teams do not need to open inbound connections for Cursor to reach the execution environment.

Read more:


World Labs Announces Atlas to Generate and Reconstruct 3D Worlds

https://www.worldlabs.ai/blog/atlas

The News:

  • World Labs announced Atlas, a multimodal world model that natively operates on text, images, and video and can generate and reconstruct 3D spaces. It generates up to one minute of camera-controlled 1440p video.
  • The model reconstructs real-world scenes into explicit 3D point clouds and Gaussian splats from as few as two or three photographs.
  • It uses a multimodal autoregressive diffusion transformer architecture that tracks 3D geometry by treating precise camera poses and depth maps as native inputs.
  • Atlas is restricted to early-access partners. Metir reports that the launch did not include a research paper or published pricing. The company’s older Marble model remains available through a public API.

My take: This is one of the most amazing visual demos I have seen in a long time. Take any photo of any scene, and Atlas will create a controlled virtual fly-through of it. This is one of those WOW moments you really have to experience to believe. If you have a few minutes left before your next meeting, check out their launch page and scroll down a bit. I guarantee this is something you have never seen before.

Read more:


Runway Solaris Generates Real-Time App Interfaces Frame by Frame Without Code

https://runway.com/news/research/introducing-solaris

The News:

  • On August 31, Runway introduced Solaris, an Interface World Model that synthesizes interactive applications and websites frame by frame in real time without generating underlying code.
  • Built on the company’s Gen-4.5 video model, Solaris renders 720p video at interactive speeds, updating the scene as users click, drag, or type.
  • In a 250-person study conducted by Runway, participants preferred Solaris over a Claude Opus 5-coded interface 71% to 21% for natural interaction.
  • Runway acknowledges significant unresolved challenges, including generating stable text, maintaining visual coherence over long sessions, and supporting accessibility tools like screen readers.
  • Runway is accepting requests for early access, and its announcement gives no public release date or details about API access or pricing.

My take: “What happens when an operating system generates apps and websites as you use them?” This is the question Runway opens with when they introduce Solaris, a world model used to create computer interfaces in real time based on your input. If you are even slightly interested in computer interface design and UX, just head over to their web page and watch the videos right now.

This is a very early prototype, and Runway itself says that “Solaris is our bet that these conceptual and technical barriers can be overcome.” Their goal is real-time interaction, coherence over an entire session, and visual quality that looks sharp and concise at 720p. This is not a replacement for the traditional computer software interface, but rather a tool you use for interactive showrooms and experiences. Instead of having to create all the 3D objects and scenes for a high-quality interactive experience, you can just take a photo or an AI-generated image, and Solaris will help you with the rest.

Read more:


OpenClaw 2.0 Launches With 16,000 Pull Requests and Multiplayer Cloud Sessions

https://openclaw.ai/blog/openclaw-2-accidentally

The News:

  • On August 31, the open-source AI agent project OpenClaw released version 2.0, its largest update ever with over 16,000 merged pull requests, adding a rebuilt browser app and shared multiplayer sessions.
  • The guided onboarding now reuses existing ChatGPT or Claude subscriptions, API keys, and local models to help users start their first conversation faster.
  • Amid ongoing vulnerabilities, the update blocks unauthenticated network exposure.
  • According to the project blog, the release was built by 933 contributors over nearly seven weeks, pausing the project’s previous pace of shipping updates every day or two.

My take: OpenClaw 2.0 is a usability release, moving from a single-user automation tool to a collaborative shared workspace where users can hand off agentic tasks without losing context. But just look at that figure - 16,000 pull requests in seven weeks (!). According to OpenClaw’s release post, this update contains roughly half of all pull requests ever merged into the project. I’m glad I am not the maintainer of OpenClaw; I can only imagine how the codebase might look under the hood.

Read more:


European Commission Designates ChatGPT a Very Large Search Engine Under the DSA

https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1772

The News:

  • On August 31, the European Commission designated OpenAI’s ChatGPT as a Very Large Online Search Engine and Reddit and Roblox as Very Large Online Platforms under the Digital Services Act.
  • The designation triggers at the 45 million average monthly EU user threshold, which all three services met; according to People of Internet, their self-reported counts were 159.1 million for ChatGPT, 57.2 million for Reddit, and 46.6 million for Roblox.
  • The companies have four months, by January 2027 according to the Commission, to comply with the strictest regulatory tier, which mandates transparency reporting and data access for vetted researchers.
  • This is the first time the EU has placed an AI chatbot in this supervision tier, where non-compliance can result in fines of up to 6% of global revenue.
  • Classifying ChatGPT as a search engine instead of a platform prevents OpenAI from using the legal safe harbor protections that platform law provides to companies hosting third-party content.

My take: This is the first time an AI chatbot has been classified as a search engine, which means OpenAI now has to set up strict governance structures and provide data access for vetted researchers to their internal platforms. In practice, I think this will have little effect on end users. OpenAI will have to address issues related to illegal content, harm to minors, physical and mental well-being, fundamental rights, elections, and public security. They are a large company, and I believe they have most of this set up already.

Read more: