CHECK

Reality Check

Official numbers cross-checked against independent measurements

Official numbers cross-checked against independent measurements

5 articles

AI & Agent

Prompt Caching (Part 2): Five Vendors, Three Clouds, One Bill

AWS documents a 4,096-token minimum for Sonnet 4.5 on Bedrock; I measured 1,024. Multipliers, TTL tiers and cross-region behaviour for five vendors — measured kept apart from documented.

AI & Agent

DeepSeek V4 Pro Goes GA: Third-Party Benchmarks Don't Match the Leaked Scorecard

DeepSeek V4 Pro 0813 launches as a GA release. Artificial Analysis gives it 53 points, ranking 23rd. The leaked agent benchmark table claims 87.9 on Terminal Bench — beating Opus 4.8 — but independent testing shows 79%, a 9-point gap that flips the result. Here's the full cross-check.

AI & Agent

Kimi K3 Is Here: 2.8 Trillion Parameters, Fully Open Source — This Time It's Different

Moonshot AI releases Kimi K3: 2.8 trillion parameters, the world's first open-source 3T-class model, 1M-token context, and native multimodality. A deep dive into the KDA and AttnRes architecture innovations, its #4 global ranking on Artificial Analysis, pricing that matches Claude Sonnet 5, and the four long-term ways an open frontier model reshapes the industry.

Tech Deep Dive

Claude Code vs OpenClaw: 510K vs 530K Lines Source Code Showdown

After Claude Code's source leak exposed 512K lines of TypeScript, we finally get a true apples-to-apples comparison with OpenClaw — architecture, agent definitions, security, and design philosophy.

Tech Deep Dive

OpenClaw vs Claude Code: Architecture and Strategy Compared

Two AI agent products, two radically different philosophies. A deep comparison of architecture, adoption, and what's next.