12 articles

Reading notes · Customer case studies · OpenAI · Added

How AI-native companies turn workflows into operating capability →

Examples from Basis, Clay, and Exa of turning recurring work into reusable agent workflows.

Reading notes · Evaluation methodology · Arena · Added

Factuality in the Arena →

Why user preference alone does not settle factual accuracy, and how Arena adds factuality signals.

Reading notes · Model evaluation · Artificial Analysis · Added

Benchmarking GPT-6 Astra →

A dated evaluation of GPT-6 Astra across capability, reasoning settings, token use, and task cost.

Reading notes · Evaluation methodology · Artificial Analysis · Added

Announcing Artificial Analysis Capability Indices v1.1 →

How domain-specific benchmark slices and weights change the meaning of a model capability score.

Reading notes · Engineering practice · Anthropic · claude.dev · Added

Lessons from building Claude Code: Prompt caching is everything →

How Claude Code organizes stable context, tools, and compaction around prompt-cache reuse.

Reading notes · Engineering practice · Anthropic · claude.dev · Added

Lessons from building Claude Code: How we use skills →

Lessons from organizing, sharing, composing, and measuring skills used inside Anthropic.

Reading notes · Engineering and usage guide · Anthropic · claude.dev · Added

The new rules of context engineering for Claude 5 generation models →

An argument for simpler project instructions and retrieving specialized context when it is needed.

Reading notes · Engineering report · Cursor · Added

How Cursor Router chooses the right model for the task →

How task complexity and model strengths inform automatic model selection in Cursor.

Reading notes · Product and workflow report · Cursor · Added

Introducing Projects →

Cursor’s approach to coordinating agents across features, migrations, and recurring maintenance.

Reading notes · Engineering report · Cursor · Added

Improved token efficiency for longer agent runs →

Where long agent runs spend tokens, and how Cursor reduces repeated context and unnecessary tool overhead.

Reading notes · Benchmark research · Arena · Added

HarnessTax: How Much Does the Harness Matter for Coding Agents? →

A comparison of model–harness combinations that puts task success and cost next to each other.

Reading notes · Usage guide · Anthropic · claude.dev · Added

Getting the most out of Opus 5.5 in Claude and Claude Code →

Practical advice for specifying a task, steering a longer run, and checking the result in Claude and Claude Code.