DISCOVER
Explore AI news and research notes.
News, practical guides, and selected reading in one place. Every update keeps its source and review context.
51 articles
Repeated document QA: does prompt order reduce cost? →
In a 24-request pilot, document-first prompts reused 12,800 input tokens. See initial versus later requests, estimated costs, answer failures and the reproducible data.
Guides & field notes · RAG evaluation · SaveMyToken · ReviewedFAQ RAG: fixed windows or paragraph chunks? →
A 14-question pilot found equal strict answer scores, with 13.0% fewer input tokens for paragraph chunks. Inspect the failures, dataset and runnable experiment.
Guides & field notes · Personal agents · SaveMyToken · ReviewedAn agent keeps interrupting you: adjust its notifications →
Trace repetitive alerts to their tasks, separate check frequency from notification rules, and verify that a quieter agent still reports what matters.
Guides & field notes · Personal agents · SaveMyToken · ReviewedCompare agent workflows: time, corrections, and cost →
Use a complete fictional inbox task, fixed acceptance criteria, and a blank results sheet to measure useful outcomes without inventing a leaderboard.
Guides & field notes · Agent workflows · SaveMyToken · ReviewedRoute an agent request with a decision model →
Define clear intent labels, keep uncertain cases visible and validate the selected tool before it runs.
Guides & field notes · Model routing · SaveMyToken · ReviewedWhen do decision models actually save money? →
Measure complete-task cost, fallback frequency and accepted quality before claiming token savings.
Guides & field notes · RAG evaluation · SaveMyToken · ReviewedCheck RAG evidence before generating an answer →
Evaluate relevance and evidence coverage after retrieval, and measure what filtering removes.
Guides & field notes · Personal agents · SaveMyToken · ReviewedYour agent remembers it wrong: fix stale facts and conflicting instructions →
Trace an incorrect answer to its source, make a scoped correction, and verify the result across new conversations and recurring tasks.
Guides & field notes · Personal agents · SaveMyToken · ReviewedMove AI memories into Claude—and check what survived →
Find the correct import flow, audit a small set of memories, test behavior in a fresh conversation, and repair omissions without assuming a complete transfer.
Guides & field notes · Decision models · SaveMyToken · ReviewedJev vs Laya vs Kev: choose a decision model →
Compare API access, local deployment and probability outputs, then build a shortlist around your task.
Guides & field notes · Agent migration · SaveMyToken · ReviewedSwitch agents and keep the settings that matter →
An agent migration guide to preserving preferences, project rules, memory, Skills, MCP connections, and automations—with a settings worksheet and checks for what actually transferred.
Guides & field notes · Personal agents · SaveMyToken · ReviewedStart with Muse: hand over your habits and project context →
Give Meta's personal agent a clear first task, a reviewed background brief, and explicit boundaries before adding connections or recurring work.
Guides & field notes · Personal agents · SaveMyToken · ReviewedBefore switching AI assistants, write your personal brief →
Carry your working preferences and active projects into a new assistant with a reviewable brief, a project card, and a small acceptance check.
Guides & field notes · Model evaluation · SaveMyToken · ReviewedHow to read model benchmarks →
Understand common AI benchmarks, score types, test settings, and the limits of a leaderboard rank.
Reading notes · Customer case studies · OpenAI · AddedHow AI-native companies turn workflows into operating capability →
Examples from Basis, Clay, and Exa of turning recurring work into reusable agent workflows.
Reading notes · Evaluation methodology · Arena · AddedFactuality in the Arena →
Why user preference alone does not settle factual accuracy, and how Arena adds factuality signals.
Reading notes · Model evaluation · Artificial Analysis · AddedBenchmarking GPT-6 Astra →
A dated evaluation of GPT-6 Astra across capability, reasoning settings, token use, and task cost.
Reading notes · Evaluation methodology · Artificial Analysis · AddedAnnouncing Artificial Analysis Capability Indices v1.1 →
How domain-specific benchmark slices and weights change the meaning of a model capability score.
Reading notes · Engineering practice · Anthropic · claude.dev · AddedLessons from building Claude Code: Prompt caching is everything →
How Claude Code organizes stable context, tools, and compaction around prompt-cache reuse.
Reading notes · Engineering practice · Anthropic · claude.dev · AddedLessons from building Claude Code: How we use skills →
Lessons from organizing, sharing, composing, and measuring skills used inside Anthropic.
Reading notes · Engineering and usage guide · Anthropic · claude.dev · AddedThe new rules of context engineering for Claude 5 generation models →
An argument for simpler project instructions and retrieving specialized context when it is needed.
Reading notes · Engineering report · Cursor · AddedHow Cursor Router chooses the right model for the task →
How task complexity and model strengths inform automatic model selection in Cursor.
Reading notes · Product and workflow report · Cursor · AddedIntroducing Projects →
Cursor’s approach to coordinating agents across features, migrations, and recurring maintenance.
Reading notes · Engineering report · Cursor · AddedImproved token efficiency for longer agent runs →
Where long agent runs spend tokens, and how Cursor reduces repeated context and unnecessary tool overhead.
Reading notes · Benchmark research · Arena · AddedHarnessTax: How Much Does the Harness Matter for Coding Agents? →
A comparison of model–harness combinations that puts task success and cost next to each other.
Reading notes · Usage guide · Anthropic · claude.dev · AddedGetting the most out of Opus 5.5 in Claude and Claude Code →
Practical advice for specifying a task, steering a longer run, and checking the result in Claude and Claude Code.
Guides & field notes · Coding plans · SaveMyToken · ReviewedCoding plans vs. API billing →
Choose a billing approach by checking compatibility, usage limits, and your actual workload.
News · Models · DeepSeek · Source checkedDeepSeek documents its current Flash API model name →
A configuration note for applications using DeepSeek's Flash model aliases.
News · Agents · DeepSeek · PublishedDeepSeek Harness previews a Claude Code Mods compatibility layer →
The October 3 alpha adds plugin creation tools, while describing Mods compatibility as an experiment.
News · Agents · Anthropic · PublishedClaude Code adds TypeScript mods for tools and interface behavior →
Mods extend the CLI and desktop app through plugins, including prompt, tool-call, and interface changes.
News · Models · Google · PublishedGoogle announces Gemini 4 Argon with a limited initial rollout →
The first access is for trusted cyber defenders through Fairwind; broader availability remains a later step.
News · Agents · DeepSeek · PublishedDeepSeek Harness Desktop bundles its plugin-management command →
The v0.2.0-rc.2 pre-release reduces setup steps on macOS and Windows and fixes desktop environment handling.
News · Costs · OpenAI · PublishedGPT-6.1 Sol arrives with a lower cached-input price →
OpenAI introduces the new Sol model for coding and professional work, with cached input listed at $0.10 per million tokens.
News · Agents · OpenAI · PublishedOpenAI adds browser computer use to the Agents API →
The September 29 changelog adds a hosted browser, with website approvals and sign-in handled by the application.
News · Models · Anthropic · PublishedClaude Sonnet 5.5 keeps token prices while targeting lower task cost →
Anthropic attributes the reported savings to token efficiency, rather than a reduction in Sonnet's standard token rates.
News · Integrations · Anthropic · PublishedClaude opens a submission portal for plugins →
Developers can submit connectors or plugin bundles, follow review feedback, and inspect usage after publication.
News · Models · OpenAI · PublishedOpenAI fixes image encoding in GPT-6 Sol and Luna →
A September 25 fix affects visual tasks in the API and Codex, including computer use.
Guides & field notes · Agent workflows · SaveMyToken · ReviewedSet a budget for an agent task →
Control retries and tool output, then verify.
Guides & field notes · Context · SaveMyToken · ReviewedKeep context useful, not endless →
Trim repeated history and oversized tool output.
Guides & field notes · Model selection · SaveMyToken · ReviewedUse the right model for each step →
Route routine work and define clear upgrade rules.
Guides & field notes · Caching · SaveMyToken · ReviewedPrompt caching, explained →
Learn what is reused and check cache usage.
Guides & field notes · Structured output · SaveMyToken · ReviewedAsk for the output you actually need →
Reduce extra generation and parsing failures.
News · Agents · DeepSeek · PublishedDeepSeek Harness adds reminder history and persistent scheduled tasks →
The v0.1.7-rc.2 pre-release adds scheduling controls and clarifies what happens when the desktop window closes.
News · Integrations · Google · PublishedGemini expands Connected Apps for project and creative work →
Google begins rolling out integrations including Linear, Airtable, Adobe, and Webflow.
News · Costs · Anthropic · PublishedClaude Opus 5.5 lowers input, output, and cache-read rates →
Anthropic announces lower standard API prices and reports efficiency gains for longer agent tasks.
News · Models · OpenAI · PublishedGPT-6 Sol and Luna expand OpenAI's lower-cost model options →
The September 22 launch adds two model tiers for coding, computer use, and everyday professional tasks.
News · Agents · Anthropic · PublishedClaude brings Cowork tasks into the regular conversation flow →
A gradual rollout to Pro and Max combines quick questions and longer tasks without a separate mode choice.
News · Costs · OpenAI · PublishedOpenAI details analytics for ChatGPT Work and Codex spending →
The Admin Console connects usage and task categories with engineering outcome indicators.
News · Agents · DeepSeek · PublishedDeepSeek Harness previews MCP resources and remote workspaces →
The September 15 alpha adds resource discovery, SSH workspace tools, and experimental browser and computer use.
News · Models · Google · PublishedGemini 3.8 Live adds a choice of voice reasoning models →
Google introduces a regular Live model and an Extended Thinking option for more complex voice workflows.
News · Integrations · Anthropic · PublishedSalesforce comes to Claude through a beta sales plugin →
The integration brings account and pipeline workflows into Claude, with beta access subject to Salesforce approval.