Tech Insights 2026 Week 34
August 17, 2026
I spend much of my time writing texts in Obsidian. But there are still many things I find annoying that require lots of manual, repetitive work.
For example, let’s say I read the news “OpenAI previews frontier speed boost” in the Rundown AI newsletter. I click on the link and save it to my notes. It gets saved as https://openai.com/index/previewing-ultrafast/?utm_source=www.therundown.ai&utm_medium=newsletter&utm_campaign=openai-feels-the-frontier-need-for-speed&_bhlid=c3e6e77f364740a88d5830a925b2061267fd1238
I then need to edit the URL manually to remove all this tracking crap. If I want to show the title of the page, I also need to convert it into a markdown link and copy the web page title by hand.
So last week I created a new plugin. This time I wanted to see how quickly I could do it with Anthropic Fable. I started last Wednesday, and the plugin is launching today on August 17. It is called Better Paste.
Now when you paste something into Obsidian, Better Paste:
- Saves remote images locally in the vault.
- Removes all tracking from URLs.
- Fetches web titles for all links automatically. So the link above would just show as Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed | OpenAI including a cleaned up link.
- Cleans up AI-generated output. It converts curly braces to plain quotes, removes hidden characters, converts en-dashes and em-dashes to hyphens, and much more.
- Transforms commas if you prefer them outside strings “like this”, instead of “like this,” with a simple setting.
As usual, the plugin comes in 21 languages, has its own web page at betterpaste.md (also in 21 languages), and it is super stable. And as usual everything is done 100% with AI. The only reason it took four days was getting the use cases and interface design right.
After this week, the journey ahead is clear to me. Software development no longer has anything to do with code structures and software architectures. This is now solved by large models like Anthropic Fable and the upcoming OpenAI GPT-6. From now on, software development is all about interface design, user testing, iteration, and saying no to a thousand things. Where most developers write code blocks with AI today, they will write complete software architectures next year.
Thank you for being a Tech Insights subscriber!
Listen to Tech Insights on Spotify: Tech Insights 2026 Week 34 on Spotify
Notable model releases last week:
- Flux TTS by Deepgram. Conversation-aware TTS model for real-time voice agents, adding cross-turn context and barge-in feedback, with cloud, VPC, and on-prem deployment and 3.4% WER on hard prompts.
- Gemini 3.7 Flash by Google. Raises FrontierCode 1.1 Main from 34.4% to 43.6% and AutomationBench from 17.0% to 30.4%, with API pricing halved through 2026.
- GPT-5.6-Cyber by OpenAI. Raises Advanced Cybersecurity Completion Rate to 95.0%, from 57.3% for GPT-5.5-Cyber, with access restricted to approved defenders through Daybreak Red.
- MAI-Code-1.1-Flash by Microsoft AI. Coding model deployed in GitHub Copilot, improving Terminal-Bench 2.1 by 22%, using 25% fewer tokens, and costing one quarter as much as 1.0.
- MAI-Thinking-1 by Microsoft AI. Sparse MoE reasoning model with 35B active and about 1T total parameters, matching Claude Opus 4.6 on SWE-Bench Pro and scoring 94.5% on AIME 2026.
- Muse Glimmer by Meta Superintelligence Labs. 30B open-weight agentic model with interleaved text-image input, Apache 2.0 licensing, 17GB 4-bit quantization, and DFlash speculative decoding delivering 3.1x faster generation on RTX 5090.
- Nemotron 3.5 Lightning by NVIDIA. Open 30B MoE model with 3B active parameters for high-volume specialized agent tasks, delivering up to 4x the output speed of similar-sized models.
- Qwen3.8-2.4T-A95B by Qwen. Open-weight 2.4T-total, 95B-active MoE text model with 262K native context and mandatory thinking mode, underpinning Qwen3.8-Max’s 86.6 on Terminal Bench 2.1.
THIS WEEK’S NEWS:
- OpenAI Reports Frontier Firms Generate 8.3x as Many Output Tokens per Active User as Typical Firms
- OpenAI Previews Ultrafast Tier Running GPT-5.6 Sol at up to 750 output tokens per second
- OpenAI Launches Computer History to Turn Mac Activity Into ChatGPT Context
- SpaceXAI Releases Grok Bot to Run Persistent AI Agents on Dedicated Cloud Computers
- Anthropic Embeds Invisible Text Watermarks in Claude Outputs to Comply With EU AI Act
- Anthropic Upgrades Claude in Chrome to Persistent Cross-Device Sessions
- Anthropic Makes Claude Code Auto Mode Default After Blocking 89% of Dangerous Commands
- DeepSeek Releases Open-Source Agent Harness
- DeepSeek Releases V4 Pro and Introduces Peak and Off-Peak API Pricing
- Lovable Raises $400 Million at $13.3 Billion Valuation
OpenAI Reports Frontier Firms Generate 8.3x as Many Output Tokens per Active User as Typical Firms
https://openai.com/index/how-enterprises-put-ai-to-work/

The News:
- OpenAI published research on enterprise AI adoption, revealing that top-tier “frontier firms”, the top 10% by AI usage each month, now generate 8.3x as many output tokens per active user as typical firms as they shift toward agentic task execution.
- As of June, Codex produced 64% of all combined Codex and ChatGPT output tokens across enterprise customers, signaling a move from chat assistance to delegated, multi-step workflows.
- Since February, weekly active enterprise Codex users grew 108x in legal and 41x in sales, compared to just 5x in software engineering.
- Early-career workers use AI more, sending 13 more messages per week than executives six months after adoption.
- Frontier firms use advanced tools more frequently, with 21% of active users utilizing Plugins weekly versus 9% at typical firms.
My take: The gap between advanced and average users is widening by the month. Ask any company out there and they have all kinds of AI initiatives, but every single one is also aware of the differences in productivity between the top 10% of AI users and the rest. This is where we focus most of our work at TokenTek right now.
It’s not so much about getting companies to start using AI; almost every large company has a dedicated group that’s pretty good at it today. The key is to keep everyone else in the company from falling behind. You do this by setting processes, rolling out tools, defining skills, and having external experts coach the teams to get everyone up to speed - not just a select few.
Read more:
- Forrester: The State of Agentic AI in 2026
- OpenAI: Enterprise Signals
- OpenAI: Running Codex Safely
- X / Inference Engine: Reaction to Codex Growth Outside Engineering
OpenAI Previews Ultrafast Tier Running GPT-5.6 Sol at up to 750 output tokens per second
https://openai.com/index/previewing-ultrafast/

The News:
- On August 13, OpenAI announced a limited preview of Ultrafast, a new API service tier powered by Cerebras chips that runs GPT-5.6 Sol at up to 750 output tokens per second.
- OpenAI says Ultrafast runs GPT-5.6 Sol up to 14 times faster than Standard processing.
- The system runs on Cerebras Wafer-Scale Engine hardware, which accelerates inference by storing model weights on-chip rather than retrieving them from external memory.
- According to HPCwire, Cerebras says the setup completed the 2,500-question Humanity’s Last Exam benchmark in just over 11 hours, compared to more than three days of continuous compute for Anthropic’s Claude Fable 5.
- The tier is currently invite-only, and OpenAI has not published pricing, a general availability date, or a dedicated API model identifier.
My take: For me personally, the token speed of Anthropic Fable 5 and GPT-5.6 Sol in fast mode is not an issue. I often run multiple sessions in parallel, and it’s very rare that I sit and get bored at the computer. So I don’t think this 14x speed multiplier is aimed at regular desktop users; it’s for specific use cases. Whether it will work for those use cases depends on price and availability, and right now we have neither. We just know that internal benchmarks show it’s possible, but not at what price.
My main question is whether it will also work with the upcoming GPT-6, which will be a much larger model. Can that one also run on the Cerebras chips? Right now I prefer Fable over GPT-5.6 for most of my use cases, and there is no way I would keep using GPT-5.6 once GPT-6 is out, even if it is 14 times faster.
Read more:
- Cerebras: Accelerating GPT-5.6 Sol Ultrafast with OpenAI
- Digital Applied: GPT-5.6 Sol Ultrafast benchmark caveats
- Hacker News: Discussion of GPT-5.6 Sol Ultrafast
- TechTimes: Ultrafast arrives without pricing or a release date
- X / Anton: A Sol and Luna multi-agent workflow
OpenAI Launches Computer History to Turn Mac Activity Into ChatGPT Context
https://learn.chatgpt.com/docs/customization/computer-history

The News:
- On August 13, OpenAI released Computer History for the ChatGPT macOS app, an opt-in feature that records daily app and website activity to build a searchable memory timeline for ChatGPT and Codex.
- The system records interaction events like clicks and typing rather than capturing screenshots or audio.
- Temporary event files are kept on the Mac for up to 48 hours, but OpenAI processes those files on its servers to generate memories.
- The feature is off by default for Pro, Business, and Enterprise users, and workspace administrators must explicitly grant access before workspace members can opt in.
- The official documentation warns that Computer History increases the risk of prompt injection from content in apps and websites.
My take: It’s easy to see how sending every single click and keyboard input to an AI model can go completely wrong. While Computer History does detect browser windows running in private mode (incognito), there are many other situations where you might type extremely sensitive information that you do not want the AI to permanently record to its memory. Hidden notes, protected vault items, temporary chats - the list just goes on.
As expected, this feature is disabled in the EU at launch. I see this as a giant experiment, and I am actually a bit surprised they pushed it straight into their main ChatGPT app. I think we will see many interesting stories about this in the coming weeks.
Read more:
- 9to5Mac: ChatGPT for Mac Adds Computer History
- Reddit: Discussion of OpenAI’s Computer History
- The Register: OpenAI Replaces Recall-Style Screenshots With “Friendly Keylogging”
- X / Michel Diz: Comparison With Microsoft Recall
SpaceXAI Releases Grok Bot to Run Persistent AI Agents on Dedicated Cloud Computers

The News:
- SpaceXAI launched Grok Bot in beta, giving AI agents access to a shared persistent cloud computer to complete multi-step tasks across apps and websites even with the user’s laptop closed.
- The agents use vision models and direct computer control to navigate graphical user interfaces, allowing them to operate software that lacks a clean API or MCP connector.
- All of a user’s Bots operate on a single shared Linux environment with unified browser cookies, files, and credentials, meaning they do not have separate security boundaries.
- The system runs on the new Grok 4.6 model, released on August 12, which, according to SpaceXAI’s own evaluation table, scored 61 on the Artificial Analysis Intelligence Index to match OpenAI’s GPT-5.6 Sol.
- Beta access is available with plans starting at $120 per seat per month for Cursor Premium Teams or $200 per month for individual Cursor Ultra plans.
My take: This is SpaceXAI’s response to Claude Cowork, ChatGPT Work, and Copilot Cowork. The main difference is that these agents run on a cloud virtual machine instead of a virtual machine on the user’s computer. I am guessing they got most of this technology from Cursor, which recently acquired other companies with similar tech. As is usual with Grok, security is not a key concern, and all agents operate in a live shared space instead of secure sandboxes. I would recommend caution before rolling out services on this platform.
Read more:
- Eesel.ai: Grok Bot beta review
- Hacker News: Discussion of Grok Bot
- SpaceXAI Docs: Approvals, security, and privacy
- SpaceXAI Docs: Use the computer and apps
- X / Diego Aguirre: Reaction to Grok Bot usage
Anthropic Embeds Invisible Text Watermarks in Claude Outputs to Comply With EU AI Act
https://www.anthropic.com/news/claude-text-watermark

The News:
- Anthropic detailed how it now embeds an invisible text watermark in outputs from all Claude models launched since August 2 to comply with the European Union AI Act.
- The system uses Google DeepMind’s SynthID-Text method to subtly bias the model’s random word choices, leaving a detectable statistical pattern without altering the meaning.
- Anthropic states the watermark adds no extra tokens, has a negligible impact on generation speed, and caused no statistically significant drop in output quality during human evaluations.
- The mark is naturally weaker on code or strict factual statements because exact answers leave no alternative word choices for the pattern to use.
- Anthropic claims the mark survives light editing.
My take: Last week Anthropic announced that all new models launched after August 2 will have an invisible watermark embedded in the text. Initially, I thought this wasn’t so bad, since you would need a special tool owned by Anthropic to detect it. Well, last week Anthropic also announced they will soon provide a “watermark detection API” so everyone can check “if a piece of text was written by Claude”. This is interesting because it might change how everyone uses AI to generate almost any text published today. If you knew that your mail recipient would get AI-generated text highlighted in their email app, would you still use AI to send it?
If you’re interested in how it works “under the hood”, [Thariq at Anthropic put together a very nice artifact](Claude Artifact) that lets you explore how digital watermarking works on different texts. The image above comes from one of his examples.
Read more:
- arXiv: Robust Text Watermarking for Large Language Models
- Forbes: Claude’s Invisible Watermarks and Professional Concerns
- Hacker News: Discussion of Claude’s Text Watermarking
- Nature: Can Anthropic’s Invisible Watermarks Curb AI Slop?
- The Next Web: Anthropic Marks Claude Outputs Worldwide
Anthropic Upgrades Claude in Chrome to Persistent Cross-Device Sessions
https://claude.com/claude-in-chrome

The News:
- Anthropic upgraded its Claude in Chrome browser extension into a persistent Claude Cowork session that syncs across devices, following an initial 1,000-user pilot.
- The extension navigates analytics dashboards, organizes Google Drive folders, and logs sales calls to the user’s CRM by reading and interacting directly with the user’s active tabs.
- Conversation history and custom skills now save to the Anthropic account, allowing a research task started in Chrome to be continued later on the desktop or mobile app.
- Anthropic explicitly warns users to avoid financial transactions and password management, noting that its built-in defenses against prompt injection attacks are not foolproof.
- According to Zenity Labs, its researchers demonstrated this vulnerability by using a hidden instruction in an email to trigger a full account takeover across a user’s authenticated Slack, X, and Claude accounts.
My take: For usefulness, this is a 10/10. You can start a task in the chat or on your mobile, open the browser to continue from there, and allow it to surf the web using your credentials. From a risk point of view, I also think it’s close to a 10/10. Using it safely with sites you trust should be no problem, but leaving it unattended to browse unknown sites could lead to disaster - especially with researchers at Zenity Labs demonstrating the theoretical possibility of complete account takeovers using instructions hidden in web pages. I think this is a great tool, but just like the OpenAI computer recorder, it needs to be used with extreme caution.
Read more:
- Anthropic: Piloting Claude in Chrome
- Claude Help Center: Claude in Chrome permissions guide
- Claude Help Center: Get started with Claude in Chrome
- X / AI Mira: Single-profile complaint
- Zenity Labs: Claude in Chrome account-takeover chain
Anthropic Makes Claude Code Auto Mode Default After Blocking 89% of Dangerous Commands
https://claude.com/blog/auto-mode-default-in-claude-code

The News:
- On August 14, Anthropic made “auto mode” the default setting for new Claude Code sessions on Pro, Max, and Team plans, using a classifier to screen tool calls instead of prompting for manual approval.
- In a study of 1,053 paid testers, human reviewers caught only 13.6% of deliberately planted dangerous commands, while the automated classifier blocked 89%.
- Anthropic found that developers manually approve 97% of permission prompts, suggesting many users click through reflexively, and that among flagged sessions those with manual approval contained serious unintended harm more than twice as often as auto mode.
- In a third-party evaluation by Trajectory Labs, zero out of 720 prompt injection attacks succeeded against Claude models running auto mode, whereas GPT-5.6 Sol had a 5.83% success rate on Codex’s Auto-review mode.
- Anthropic is no longer charging Pro, Max, and Team users for the extra tokens consumed by the auto mode classifier, though the feature remains opt-in for Enterprise and API accounts.
My take: “Users approve 97% of permission prompts in Claude Code”. Based on that, having users approve anything an AI asks for is close to useless. They will just keep spamming the approve button. If you are not running your AI agents in a sandbox I think this new auto mode is a great thing. I just wish Microsoft and OpenAI had something similar for Copilot and Codex.
Read more:
- arXiv: A Stress-Test Evaluation of Claude Code’s Auto Mode
- Claude Code Docs: Configure Auto Mode
- Simon Willison: Reaction to Claude Code Auto Mode
- X / Jyothi Venkat: Token Cost Concern About Auto Mode
- X / Prasenjit Sarkar: Agent Fleet Governance Concern
DeepSeek Releases Open-Source Agent Harness
https://github.com/deepseek-ai/deepseek-harness

The News:
- DeepSeek released DeepSeek Harness, an MIT-licensed developer preview of an open-source agent harness that, according to AI/TLDR, passed 34,000 GitHub stars on its first day.
- The framework uses an architecture where “everything is a plugin”, allowing developers to mix, match, and replace the models, tools, storage, sandboxes, and user interface.
- It installs as an npm package and serves a local web application.
- According to AIReiter, DeepSeek used a minimal version of this harness to generate its official V4-Flash agent benchmark scores on July 31.
- According to AIReiter, an unrelated Python library called
deepseek-harnesshas existed on PyPI since May, but the official DeepSeek release is a Node.js project.
My take: So, what can you do with this? With 34,000 GitHub stars the first day, the answer should be a lot of things. In my experience, most use cases that involve AI can often be solved with straightforward workflows, where LLMs are called for specific tasks with deterministic output. If you need slightly more autonomy, you wire up an agentic system like Agno or LangGraph with your own services, databases, FastAPI backend, etc. Now let’s say you want to build your own “agent computer” with a shell, files, codebase, and session UI - then you would use DeepSeek Harness. You can see it as three different levels of agentic systems. I am just surprised that the interest is so high for something this complex.
Read more:
- AIReiter: DeepSeek Harness installation and package warning
- DeepSeek: Harness developer preview
- Hacker News: Discussion of DeepSeek Harness developer preview
- MindStudio: Hands-on test of DeepSeek Harness
- The Register: Why the agent harness matters
DeepSeek Releases V4 Pro and Introduces Peak and Off-Peak API Pricing
https://x.com/deepseek_ai/status/2087864585504305397

The News:
- On August 13, DeepSeek moved its DeepSeek-V4-Pro model out of preview, adding adjustable reasoning levels and native OpenAI Responses API support optimized for Codex.
- Developers can now toggle the model’s reasoning effort between low, high, and max to match the complexity of their specific workflows.
- The update introduces peak and off-peak API pricing starting August 16, and deepseekv4pro.com reports the model’s cache-hit input cost will rise by 12 times during peak hours.
- Despite the launch, ReveloHQ wrote on X that the DeepSeek-V4-Flash model outscored V4-Pro by 10.8 points on Terminus 2.
My take: There are two things I found very interesting in this announcement. First, DeepSeek is no longer the cheapest tool on the block. At $1.32 per 1 million input tokens and $3.96 per 1 million output tokens, it is still cheaper than GPT and Claude, but the gap is closing. The second interesting bit here is the off-peak pricing, which is half the cost. Will this change when people work during the day? I know I would get up two hours earlier if it meant double the token usage. I will be following this closely.
Read more:
- DeepSeek: V4 Pro GA release
- FAR AI: Security stress test of DeepSeek V4 Pro
- Hacker News: Discussion of DeepSeek V4 Pro 0813
- VentureBeat: DeepSeek Harness and V4 Pro launch
- X / ReveloHQ: Reaction to V4 Flash versus V4 Pro
Lovable Raises $400 Million at $13.3 Billion Valuation

The News:
- On August 12, vibe-coding startup Lovable raised a $400 million Series C round led by Menlo Ventures and the Scaleup Europe Fund, doubling its valuation to $13.3 billion.
- The company reached $500 million in annualized run rate revenue in June and, according to Value Add VC, employees at roughly two-thirds of the Fortune 500 use Lovable, including Nvidia and Adidas.
- Users have built 60 million projects on the platform, with Lovable-generated apps now attracting 900 million monthly visitors.
- Starting September 9, Lovable will default to using customer data from Free and Pro tiers for AI model training unless users manually opt out.
My take: Lovable, the 2-year-old company, is now valued at $13.3 billion. 😳 Let that sink in for a while. In 2008 in Zimbabwe, you had to use a wheelbarrow to carry the money you needed to buy a loaf of bread, which would cost billions of Zimbabwean dollars due to inflation. You kind of get the same feeling here. It’s like money has lost any sense of realistic value for some people and companies. I love Lovable, but there is no sensible way the company is worth this amount of money.
This entire market situation is not sustainable, and financial corrections are bound to happen sooner or later. It’s just a question of how long the major companies can push it forward so they have time to collect their profits and invest in something more solid before the bubble bursts and everyone still remaining in the stock market loses all their virtual cash.
Read more: