Today’s radar was packed with model launches, but the item that stopped me was an essay about measurement. I will start with it, then share the rest of what passed through the filter.
In brief
There is a gap between the AI productivity gains teams feel and the gains that show up in metrics. Writing less, reviewing faster and escaping the blank page in seconds can speed up one step, but they do not guarantee delivered value. To decide whether AI is working, measure the end-to-end workflow before promising results: time, review, rework, integration and the business outcome that matters.
In my roadmap meetings, the phrase that makes me most suspicious is not “this is hard.” It is “everyone here felt faster.”
Today I read an essay about what people are calling the AI productivity gap. The idea is simple and uncomfortable: there is a large distance between the gain teams think they are getting from AI and the gain that actually appears in the metrics. The discussion caught fire among developers, but this is a management problem, not a code problem.
And it makes sense. The feeling of speed is real. You write less, review faster and get past the blank page in seconds. But feeling faster is not the same as delivering value. Between the two lie review, rework, integration and fixing what slipped through.
Working on credit products taught me to respect that difference in a somewhat hard way. Automating document generation is easy to demonstrate. Reducing the actual time between an originator uploading an invoice and the money reaching their account is another story, because it involves workflow, exceptions, checks and people. The first makes a good demo. The second produces results.
That difference is central to AI product management. An activity metric can show that the team produced more; an end-to-end metric shows whether there was less rework, less cycle time or a better outcome for the customer. The product management guide helps separate activity from impact, while AI for Product Managers explains how to establish a baseline before comparing changes.
My practical takeaway is a boring but useful rule: measure before you promise. Choose the end-to-end metric, see where it stands today and only then put AI into the workflow. Without that, the business case becomes a discussion of perception, and perception does not survive the first budget committee. When the initiative involves AI agents, an evaluation template helps make quality, cost, latency and oversight criteria explicit.
None of this makes me less optimistic. I think the gap signals a stage, not a failure. Every technology that changes how people work goes through a period when we feel the effect before we know how to measure it. Whoever learns to measure first will defend investment much more confidently than those who can only tell stories about tenfold gains.
If you would like to read the full essay, it is worth the time: The AI Productivity Gap.
The rest of the radar
Qwen3.8-Max with open weights — An open frontier model changes the build-versus-buy calculation and the cost ceiling for agentic features. Read more
Dynamic Workflows in Claude Code — Moves the unit of work from the prompt to the process, changing how an AI feature is specified and versioned. Read more
OpenAI cuts GPT-5.6 API pricing — An 80% drop at the lowest tier reopens use cases that were outside the margin and puts pressure on your product’s pricing. Read more
Noisy LLM evaluators are still useful — Removes the excuse that “our eval is not reliable” and makes it feasible to measure agent quality without a human gold set. Read more
Robinhood lets AI agents trade — A boundary case of an agent with write permission in a regulated domain, bringing the trust and accountability debate every agentic product will face. Read more
Algolia launches Agent Studio in beta — An agent grounded in trustworthy proprietary data becomes an off-the-shelf product rather than an internal project. Read more
Study maps dark patterns in chatbots — Defines what already counts as a manipulative pattern in conversational UX, with direct reputational and regulatory risk for the backlog. Read more
Step 3.7 Flash, from StepFun — Another fast, low-cost model puts pressure on the low-cost tier and expands the supplier set. Read more
AISlop, a CLI for AI code smells — Signals a new category of quality tooling for LLM-generated code, and a new cost in the delivery cycle. Read more
That is it for today. There will be more tomorrow.