Today the radar is heavy on costs and agents. GPT-5.6 arrived in three versions in one move, with Grok and Meta launching on the same day, and the result is another drop in the price of running AI. I cover what this changes for people deciding the roadmap and then everything else that passed the filter.

In brief

With three model options and more competing launches, inference costs are falling again. For product teams, the decision becomes which use cases now fit the budget — and how to combine models by task without giving up judgment, permissions, and governance.

Anyone who works in product knows that the line item that shows up most often in the spreadsheet is not the beautiful feature. It is the cost of running it at scale.

That is why I stopped to look at GPT-5.6 this week. OpenAI released three versions at once: a top-tier option for heavy reasoning, a middle option with quality similar to the previous generation at half the cost, and a fast, inexpensive option for high volume. It did not arrive alone: Grok and Meta launched on the same day. The practical result is that inference costs are falling again.

For people working in product — and I work on the product side of credit — this changes the decision math. An idea that did not make financial sense six months ago because of processing costs can now fit the budget. What was too expensive to put into production becomes viable.

In receivables and structured credit, this is concrete. Reading and checking documents, cross-referencing invoice information, and supporting analysis at volume are tasks that have always existed, but they depended on the automation economics working. When the cost per operation falls, what used to be a pilot becomes routine.

The point I try not to lose sight of is that a lower price is not an excuse to take judgment off the table. In finance, a poorly designed low-cost solution becomes expensive. The value lies in using the right model for each task: the expensive one only where it pays for itself, the cheaper one where volume leads. Designing AI product features around that routing is now almost a product decision.

That choice also belongs in the design of AI agents. Each stage needs a model that fits its risk, volume, and type of decision. When financial operations involve consequential actions, cost is not the only criterion: permissions and AI governance are part of the architecture.

What excites me is that this cost decline brings more people closer to the technology. Smaller teams, smaller companies, and projects that did not have enough budget room can all start testing. That is where uses nobody anticipated begin to emerge.

For anyone curious about the numbers behind this new group of models, the details are here.

The rest of the radar

OpenAI Presence — a template for operating agents in production with governance: guardrails, evaluations, and approved actions. Read more

Microsoft adds Sales and Service Agent to Outlook and Teams — agents embedded where work already happens raise the bar for adoption and distribution. Read more

Dynamic Workflows in Claude Code — dynamic workflows change how teams orchestrate coding agents and internal automations. Read more

StepFun releases Step 3.7 Flash — another fast, inexpensive model expands vendor options and cost pressure. Read more

Screenpipe (YC S26) — continuous user context is a powerful input, and a privacy risk, for agentic features. Read more

Robinhood lets AI agents trade stocks — a real case of giving agents real-world actions, with implications for trust and accountability. Read more

Noisy LLM evaluators still help — do not wait for a perfect evaluation to start measuring agent quality. Read more

OneCLI — a credential gateway that keeps secrets outside agents; security is a blocker for enterprise adoption. Read more

Promptloop — creating, running, and improving prompt evaluations from the terminal reduces regressions in AI features. Read more


That is what I set aside today. If any of these points changes how you think about the next quarter, the read was worthwhile.