DISCOVER
Explore AI news and research notes.
News, practical guides, and selected reading in one place. Every update keeps its source and review context.
12 articles
How AI-native companies turn workflows into operating capability →
Examples from Basis, Clay, and Exa of turning recurring work into reusable agent workflows.
Reading notes · Evaluation methodology · Arena · AddedFactuality in the Arena →
Why user preference alone does not settle factual accuracy, and how Arena adds factuality signals.
Reading notes · Model evaluation · Artificial Analysis · AddedBenchmarking GPT-6 Astra →
A dated evaluation of GPT-6 Astra across capability, reasoning settings, token use, and task cost.
Reading notes · Evaluation methodology · Artificial Analysis · AddedAnnouncing Artificial Analysis Capability Indices v1.1 →
How domain-specific benchmark slices and weights change the meaning of a model capability score.
Reading notes · Engineering practice · Anthropic · claude.dev · AddedLessons from building Claude Code: Prompt caching is everything →
How Claude Code organizes stable context, tools, and compaction around prompt-cache reuse.
Reading notes · Engineering practice · Anthropic · claude.dev · AddedLessons from building Claude Code: How we use skills →
Lessons from organizing, sharing, composing, and measuring skills used inside Anthropic.
Reading notes · Engineering and usage guide · Anthropic · claude.dev · AddedThe new rules of context engineering for Claude 5 generation models →
An argument for simpler project instructions and retrieving specialized context when it is needed.
Reading notes · Engineering report · Cursor · AddedHow Cursor Router chooses the right model for the task →
How task complexity and model strengths inform automatic model selection in Cursor.
Reading notes · Product and workflow report · Cursor · AddedIntroducing Projects →
Cursor’s approach to coordinating agents across features, migrations, and recurring maintenance.
Reading notes · Engineering report · Cursor · AddedImproved token efficiency for longer agent runs →
Where long agent runs spend tokens, and how Cursor reduces repeated context and unnecessary tool overhead.
Reading notes · Benchmark research · Arena · AddedHarnessTax: How Much Does the Harness Matter for Coding Agents? →
A comparison of model–harness combinations that puts task success and cost next to each other.
Reading notes · Usage guide · Anthropic · claude.dev · AddedGetting the most out of Opus 5.5 in Claude and Claude Code →
Practical advice for specifying a task, steering a longer run, and checking the result in Claude and Claude Code.