Today’s radar was packed with major model launches — GPT-5.6, ChatGPT Work, and Muse Spark 1.1. But what really caught my attention was a much quieter benchmark about accounting.

I came across a benchmark this week that made me stop and think about where we are with AI applied to finance.

An open model, GLM 5.2, was tested on a very specific task: classifying VAT entries, the kind of structured work any accountant knows by heart. The result came close to the accuracy of a human professional.

This is not a general-purpose model merely “talking well” about accounting. It is accuracy on a structured financial task — the kind that underpins entire back-office processes.

This matters directly to me because I work on the product side of credit, and much of what moves receivables, invoices, and structured transactions is also made up of tasks like these: repetitive, rules-based, and supported by data clean enough to automate for real.

For a long time, the product narrative was “AI assisting the analyst.” A result like this points to another phase: entire tasks moving to high-confidence automation, with human oversight focused primarily on exceptions.

That changes product design. It changes the level of control and auditability we need to build alongside it. And it changes the expectations of the people operating the process every day.

I do not think this replaces human judgment in more complex credit decisions. But in structured processes, the bar for what can already be automated has shifted again — and quickly.

For anyone curious to see the full benchmark, here is the link: Read more

The rest of the radar

GPT-5.6 — resets the baseline for capability and token cost that PMs use when deciding which model to embed in a product. Read more

ChatGPT Work — raises the bar for what to expect from a productivity agent that acts across connected apps rather than simply answering questions. Read more

Muse Spark 1.1 (Meta) — another API model option to evaluate before switching vendors in AI products. Read more

Zoom acquires Common Room — an M&A example of bringing AI agents into an existing product platform without building them from scratch. Read more

Context.dev (YC S26) — provides web-data infrastructure for agents and RAG, reducing the time required to build AI features. Read more

Frugon — a practical tool for controlling inference costs as AI and agent usage scales in a product. Read more

Lucid — observability and control over model reasoning, increasingly critical to the quality and predictability of AI products. Read more

Radware (Agentic AI Protection) — security and governance for autonomous agents are becoming product requirements, not optional differentiators. Read more

AI on LinkedIn (Pangram) — concrete data on the saturation of AI-generated content helps inform detection, trust, and differentiation in social and content products. Read more


That is what caught my attention today. See you tomorrow.