01 / Measure the work you still have to do
Choose something you do every week, ask two candidate assistants to do it, and record your own involvement. A fast first answer may still need extensive correction. A slower answer may leave you with less work.
This is an original comparison method, versioned October 5, 2026. It includes a complete fictional task and empty recording tables. We have not run a controlled comparison of Muse, Claude, or other products for this article. There are no measured completion times, savings percentages, or winners here.
02 / Define the deliverable before choosing the winner
The example task is to turn a few project messages into current arrangements, personal actions, and a short reply draft. It checks extraction, handling of changed information, unknowns, and writing.
Start by supplying the same material manually to both candidates. If your real need is automatic inbox access and cross-app execution, compare that complete workflow separately, including connection setup, authorization, approval waits, and result inspection.
Those are different tests. Results from pasted messages do not establish that a product can independently manage your inbox. Name the endpoint precisely: a reviewed draft, a calendar proposal, or a verified external action.
03 / Use this identical input packet
All names and messages are fictional. Assume it is Monday, October 12, 2026, and use Asia/Shanghai throughout. Give each candidate all six messages unchanged. Keep the acceptance answers below out of its input.
If you translate this packet or change the requested output language, treat it as a separate test condition. The Chinese companion uses a Chinese-character limit; its results should not be pooled with this English-word-limit task.
M1 | Monday 09:00 | Client
Let's schedule the Orange Leaf review for Tuesday at 14:00.
Please reserve 30 minutes.
M2 | Monday 09:20 | Client
Update: Move the review to Tuesday at 16:00, still 30 minutes.
This message replaces the earlier time.
M3 | Monday 09:30 | Project lead
Please deliver the six-page proposal by Tuesday at 12:00 for the review.
M4 | Monday 10:00 | Client
We will provide final feedback on Friday. The time is not yet decided.
M5 | Monday 10:10 | Events newsletter
This week's public livestream schedule is available.
Take a look if interested; no reply needed.
M6 | Monday 11:00 | Project lead
I received pages 1–4. Please add pages 5–6; the deadline is unchanged.04 / Send the same request to each candidate
Choose candidates that can process the same packet. Use the same output requirements and action scope. Do not quietly give one product extra hints or additional retries because you prefer its style.
Using only M1–M6, produce:
1. Current project arrangements: item, owner, date/time, source message ID.
2. My next actions, ordered by known deadline.
3. A reply draft to the sender of M6, at most 60 English words.
Today is 2026-10-12. Time zone: Asia/Shanghai.
Write "not determined" for information or times that were not supplied.
Return plans and a draft only. Do not send messages, create reminders,
or modify a calendar.05 / Lock the acceptance criteria before starting
Check these required conditions individually. You can also record readability, but do not let attractive phrasing compensate for a wrong time. No arbitrary combined score is needed to see which facts failed.
The table contains the expected answers for this fictional task. It is not an observed product result.
| Check | Required outcome |
|---|---|
| Latest review time | October 13, 16:00; 30 minutes; supported by M2 |
| Proposal deadline | October 13, 12:00; M3, unchanged by M6 |
| Remaining work | Complete pages 5–6; do not restart all six pages |
| Final feedback | Client owns it; October 16; exact time not determined |
| Unrelated message | M5 is not turned into a required project action |
| Reply | Acknowledges the two missing pages, does not claim completion, meets length limit |
| Action scope | No message sent, reminder created, or calendar changed |
06 / Record configuration, then four kinds of time
Before each run, record the product, plan, client, date, visible model name, memory state, and tool permissions. If a model version is not disclosed, write ‘not disclosed.’ Do not infer it from the brand name.
Record one-time setup separately. From task submission, measure time to the first returned result and elapsed time until every required condition passes. Also record your active minutes reading, correcting, editing, and checking.
Elapsed completion time already includes waiting and human participation. Do not add active minutes to it again as a total. If you can do other work while waiting, keep that distinction visible. If the task never passes, its completion time is unknown; retain elapsed time at the stopping point separately.
07 / Count correction rounds consistently
Count one correction round each time you send a message asking for a fix. Three issues in one message still count as one round. Record required action approvals separately, and keep time spent manually editing the output in your active-time record.
Give each correction a reason: ignored M2, assigned the client's feedback to me, or exceeded the reply limit. Specific reasons make the comparison useful for deciding what to adjust.
Set a stopping rule before the run, such as at most two correction rounds or ten minutes. These are optional test limits, not product capability claims. If the task still fails, record it as unfinished. Do not present the last reply time as a successful completion time.
08 / Fill in a results sheet without turning unknowns into zero
Use actual observations for every cell. If a subscription does not expose a per-task charge, record ‘subscribed; per-task cost not visible.’ That does not make the task free. For usage-based billing, use the run's actual usage and applicable rates, including retries and any separately visible charges.
Keep original outputs and a short event log. A later correction to your evaluation is much easier when you can inspect what happened.
| Field | Agent A | Agent B |
|---|---|---|
| Product, plan, client, date | To record | To record |
| Visible model, memory, tools | To record | To record |
| One-time setup | To record | To record |
| Time to first result | To record | To record |
| Elapsed time to acceptance | Unknown | Unknown |
| Active human time | To record | To record |
| Correction rounds and reasons | To record | To record |
| Approval count | To record | To record |
| All required checks passed | Not tested | Not tested |
| Confirmed incremental cost | Unknown | Unknown |
| Original output and log | To record | To record |
09 / Repeat without crediting the product for your practice
One run can reveal a useful problem but does not establish a lasting winner. For a personal comparison, try three different packets of similar difficulty and rotate which candidate goes first. Set the acceptance answers for each packet in advance.
For first-use comparisons, record existing memory and make starting conditions as comparable as practical. For a tuned-workflow comparison, give each candidate the same adjustment opportunity, then evaluate on new material. Report those two conditions separately.
For connected workflows, use equivalent test inboxes or file sets and a precise endpoint. A product without the necessary connector may be unsuitable for that workflow; there is no need to invent a failed completion time for it.
10 / Choose for your workload and preserve the evidence
First exclude options that cannot reliably meet your essential conditions. Then compare active human time, cost visibility, and recurring maintenance. First-response speed alone does not tell you how much work remains.
To estimate weekly involvement, multiply typical active minutes per task by tasks per week, then add weekly maintenance. Show one-time setup separately. This is a planning estimate for your workload, not a promise of long-term savings.
Write a conclusion with a boundary: ‘On these inputs, settings, and dates, A needed fewer corrections; B returned its first answer sooner.’ Keep failures and original results so you can repeat the comparison after your work or the products change.
