51 articles

Guides & field notes · Caching · SaveMyToken · Reviewed

Repeated document QA: does prompt order reduce cost? →

In a 24-request pilot, document-first prompts reused 12,800 input tokens. See initial versus later requests, estimated costs, answer failures and the reproducible data.

Guides & field notes · RAG evaluation · SaveMyToken · Reviewed

FAQ RAG: fixed windows or paragraph chunks? →

A 14-question pilot found equal strict answer scores, with 13.0% fewer input tokens for paragraph chunks. Inspect the failures, dataset and runnable experiment.

Guides & field notes · Personal agents · SaveMyToken · Reviewed

An agent keeps interrupting you: adjust its notifications →

Trace repetitive alerts to their tasks, separate check frequency from notification rules, and verify that a quieter agent still reports what matters.

Guides & field notes · Personal agents · SaveMyToken · Reviewed

Compare agent workflows: time, corrections, and cost →

Use a complete fictional inbox task, fixed acceptance criteria, and a blank results sheet to measure useful outcomes without inventing a leaderboard.

Guides & field notes · Agent workflows · SaveMyToken · Reviewed

Route an agent request with a decision model →

Define clear intent labels, keep uncertain cases visible and validate the selected tool before it runs.

Guides & field notes · Model routing · SaveMyToken · Reviewed

When do decision models actually save money? →

Measure complete-task cost, fallback frequency and accepted quality before claiming token savings.

Guides & field notes · RAG evaluation · SaveMyToken · Reviewed

Check RAG evidence before generating an answer →

Evaluate relevance and evidence coverage after retrieval, and measure what filtering removes.

Guides & field notes · Personal agents · SaveMyToken · Reviewed

Your agent remembers it wrong: fix stale facts and conflicting instructions →

Trace an incorrect answer to its source, make a scoped correction, and verify the result across new conversations and recurring tasks.

Guides & field notes · Personal agents · SaveMyToken · Reviewed

Move AI memories into Claude—and check what survived →

Find the correct import flow, audit a small set of memories, test behavior in a fresh conversation, and repair omissions without assuming a complete transfer.

Guides & field notes · Decision models · SaveMyToken · Reviewed

Jev vs Laya vs Kev: choose a decision model →

Compare API access, local deployment and probability outputs, then build a shortlist around your task.

Guides & field notes · Agent migration · SaveMyToken · Reviewed

Switch agents and keep the settings that matter →

An agent migration guide to preserving preferences, project rules, memory, Skills, MCP connections, and automations—with a settings worksheet and checks for what actually transferred.

Guides & field notes · Personal agents · SaveMyToken · Reviewed

Start with Muse: hand over your habits and project context →

Give Meta's personal agent a clear first task, a reviewed background brief, and explicit boundaries before adding connections or recurring work.

Guides & field notes · Personal agents · SaveMyToken · Reviewed

Before switching AI assistants, write your personal brief →

Carry your working preferences and active projects into a new assistant with a reviewable brief, a project card, and a small acceptance check.

Guides & field notes · Model evaluation · SaveMyToken · Reviewed

How to read model benchmarks →

Understand common AI benchmarks, score types, test settings, and the limits of a leaderboard rank.

Reading notes · Customer case studies · OpenAI · Added

How AI-native companies turn workflows into operating capability →

Examples from Basis, Clay, and Exa of turning recurring work into reusable agent workflows.

Reading notes · Evaluation methodology · Arena · Added

Factuality in the Arena →

Why user preference alone does not settle factual accuracy, and how Arena adds factuality signals.

Reading notes · Model evaluation · Artificial Analysis · Added

Benchmarking GPT-6 Astra →

A dated evaluation of GPT-6 Astra across capability, reasoning settings, token use, and task cost.

Reading notes · Evaluation methodology · Artificial Analysis · Added

Announcing Artificial Analysis Capability Indices v1.1 →

How domain-specific benchmark slices and weights change the meaning of a model capability score.

Reading notes · Engineering practice · Anthropic · claude.dev · Added

Lessons from building Claude Code: Prompt caching is everything →

How Claude Code organizes stable context, tools, and compaction around prompt-cache reuse.

Reading notes · Engineering practice · Anthropic · claude.dev · Added

Lessons from building Claude Code: How we use skills →

Lessons from organizing, sharing, composing, and measuring skills used inside Anthropic.

Reading notes · Engineering and usage guide · Anthropic · claude.dev · Added

The new rules of context engineering for Claude 5 generation models →

An argument for simpler project instructions and retrieving specialized context when it is needed.

Reading notes · Engineering report · Cursor · Added

How Cursor Router chooses the right model for the task →

How task complexity and model strengths inform automatic model selection in Cursor.

Reading notes · Product and workflow report · Cursor · Added

Introducing Projects →

Cursor’s approach to coordinating agents across features, migrations, and recurring maintenance.

Reading notes · Engineering report · Cursor · Added

Improved token efficiency for longer agent runs →

Where long agent runs spend tokens, and how Cursor reduces repeated context and unnecessary tool overhead.

Reading notes · Benchmark research · Arena · Added

HarnessTax: How Much Does the Harness Matter for Coding Agents? →

A comparison of model–harness combinations that puts task success and cost next to each other.

Reading notes · Usage guide · Anthropic · claude.dev · Added

Getting the most out of Opus 5.5 in Claude and Claude Code →

Practical advice for specifying a task, steering a longer run, and checking the result in Claude and Claude Code.

Guides & field notes · Coding plans · SaveMyToken · Reviewed

Coding plans vs. API billing →

Choose a billing approach by checking compatibility, usage limits, and your actual workload.

News · Models · DeepSeek · Source checked

DeepSeek documents its current Flash API model name →

A configuration note for applications using DeepSeek's Flash model aliases.

News · Agents · DeepSeek · Published

DeepSeek Harness previews a Claude Code Mods compatibility layer →

The October 3 alpha adds plugin creation tools, while describing Mods compatibility as an experiment.

News · Agents · Anthropic · Published

Claude Code adds TypeScript mods for tools and interface behavior →

Mods extend the CLI and desktop app through plugins, including prompt, tool-call, and interface changes.

News · Models · Google · Published

Google announces Gemini 4 Argon with a limited initial rollout →

The first access is for trusted cyber defenders through Fairwind; broader availability remains a later step.

News · Agents · DeepSeek · Published

DeepSeek Harness Desktop bundles its plugin-management command →

The v0.2.0-rc.2 pre-release reduces setup steps on macOS and Windows and fixes desktop environment handling.

News · Costs · OpenAI · Published

GPT-6.1 Sol arrives with a lower cached-input price →

OpenAI introduces the new Sol model for coding and professional work, with cached input listed at $0.10 per million tokens.

News · Agents · OpenAI · Published

OpenAI adds browser computer use to the Agents API →

The September 29 changelog adds a hosted browser, with website approvals and sign-in handled by the application.

News · Models · Anthropic · Published

Claude Sonnet 5.5 keeps token prices while targeting lower task cost →

Anthropic attributes the reported savings to token efficiency, rather than a reduction in Sonnet's standard token rates.

News · Integrations · Anthropic · Published

Claude opens a submission portal for plugins →

Developers can submit connectors or plugin bundles, follow review feedback, and inspect usage after publication.

News · Models · OpenAI · Published

OpenAI fixes image encoding in GPT-6 Sol and Luna →

A September 25 fix affects visual tasks in the API and Codex, including computer use.

Guides & field notes · Agent workflows · SaveMyToken · Reviewed

Set a budget for an agent task →

Control retries and tool output, then verify.

Guides & field notes · Context · SaveMyToken · Reviewed

Keep context useful, not endless →

Trim repeated history and oversized tool output.

Guides & field notes · Model selection · SaveMyToken · Reviewed

Use the right model for each step →

Route routine work and define clear upgrade rules.

Guides & field notes · Caching · SaveMyToken · Reviewed

Prompt caching, explained →

Learn what is reused and check cache usage.

Guides & field notes · Structured output · SaveMyToken · Reviewed

Ask for the output you actually need →

Reduce extra generation and parsing failures.

News · Agents · DeepSeek · Published

DeepSeek Harness adds reminder history and persistent scheduled tasks →

The v0.1.7-rc.2 pre-release adds scheduling controls and clarifies what happens when the desktop window closes.

News · Integrations · Google · Published

Gemini expands Connected Apps for project and creative work →

Google begins rolling out integrations including Linear, Airtable, Adobe, and Webflow.

News · Costs · Anthropic · Published

Claude Opus 5.5 lowers input, output, and cache-read rates →

Anthropic announces lower standard API prices and reports efficiency gains for longer agent tasks.

News · Models · OpenAI · Published

GPT-6 Sol and Luna expand OpenAI's lower-cost model options →

The September 22 launch adds two model tiers for coding, computer use, and everyday professional tasks.

News · Agents · Anthropic · Published

Claude brings Cowork tasks into the regular conversation flow →

A gradual rollout to Pro and Max combines quick questions and longer tasks without a separate mode choice.

News · Costs · OpenAI · Published

OpenAI details analytics for ChatGPT Work and Codex spending →

The Admin Console connects usage and task categories with engineering outcome indicators.

News · Agents · DeepSeek · Published

DeepSeek Harness previews MCP resources and remote workspaces →

The September 15 alpha adds resource discovery, SSH workspace tools, and experimental browser and computer use.

News · Models · Google · Published

Gemini 3.8 Live adds a choice of voice reasoning models →

Google introduces a regular Live model and an Extended Thinking option for more complex voice workflows.

News · Integrations · Anthropic · Published

Salesforce comes to Claude through a beta sales plugin →

The integration brings account and pipeline workflows into Claude, with beta access subject to Salesforce approval.