Trust and quality notes
- Last updated
- October 4, 2026
The risky part of a new AI workflow is when a draft becomes an action someone else must live with: an email reaches a prospect, a record changes, or a promise enters a customer conversation.
Ask, “What is the least authority it needs to help us learn?” A workflow can observe and draft while a person controls the consequential step. The goal is earned trust based on real runs, visible evidence, and clear limits.
Use a trust ladder, not an autonomy switch
Treat authority as a set of levels. Each level should have its own permissions, evidence requirement, and way back down.
Level 0: Observe
The workflow reads approved inputs and produces no external output. It might classify inquiries or identify missing details. A person compares its suggestions with what happened.
Permissions: read-only access to the smallest necessary inbox, form, or record set. No write access, sending, deletion, or exporting.
Level 1: Draft
The workflow creates a draft in a designated queue. It can suggest an email, summary, or record update, but cannot apply or send it.
Permissions: read approved sources and write only to a draft location. Every draft should be visibly labeled and linked to its source inputs.
Level 2: Prepare with required approval
The workflow may fill fields or prepare an action. A named person must approve the exact output before anything consequential happens.
Permissions: narrowly scoped write access, with sending, publishing, spending, deleting, and commitments still blocked. In Agentic Workers, email sends remain approval-gated. A prepared outbound message is not permission to send it.
Level 3: Take bounded, reversible action
Only well-understood actions earn this level, such as adding an internal tag or creating a review task. Each action must be reversible and constrained by explicit rules.
Permissions: allowlisted actions, records, hours, and volume. Keep sensitive systems and irreversible actions outside the boundary.
Level 4: Expand one boundary at a time
Increase one dimension, such as volume or one record type, while every other limit stays steady. This makes the effect observable and rollback practical.
This level is not a destination every workflow needs. Drafting with approval may remain the right design for work involving reputation, judgment, private information, or commitments.
Set promotion criteria before the test
Do not promote a workflow because several outputs “looked good.” Define a review window and measurable conditions first.
For each level, record:
- the minimum number and variety of reviewed runs;
- the acceptable rate of factual corrections and policy violations;
- which mistakes are disqualifying regardless of frequency;
- who can approve promotion;
- the next level’s exact additional permission;
- the date of the next review.
Useful measures include approval rate, edits between draft and final, factual corrections, missing context, policy exceptions, duplicate actions, response time, and reversals. Segment results by inquiry type. Ten simple requests do not prove readiness for an unusual complaint.
Promotion should require clean performance within the proposed permission plus a failure test: remove a required input, introduce ambiguity, or make a destination unavailable. The workflow should route to review rather than improvise.
Define rollback and stop conditions
Trust can go down as well as up. Write the downgrade rules while everyone is calm.
Stop the workflow immediately if it sends or changes something outside scope, exposes private data, invents a fact, makes an unauthorized promise, repeats an action, loses the link between input and output, or cannot record what it did. Also stop when a connected system changes in a way that invalidates the tested process.
A rise in edits, repeated misclassification, or growing exceptions can trigger a rollback by one level. Pause actions, preserve evidence, revoke the added permission, and return affected items to review. Resume only after a documented fix and fresh bounded test.
Keep evidence without keeping everything
An audit trail should answer five questions: What input was used? What rule or instruction applied? What output was produced? Who approved it? What action occurred, and when?
Store record identifiers, timestamps, workflow version, approval status, action result, and errors. A teammate should be able to reconstruct a decision without guessing.
Privacy needs its own boundary. Collect only necessary data and restrict access by role. Do not copy full inboxes or customer records into logs when a reference and limited excerpt will do. Set retention periods for sources, drafts, approvals, and traces, then delete or redact them on schedule. Legal or customer requirements may require different periods.
A lead-response example
Imagine a small services team receives a website inquiry:
We need help cleaning up weekly sales reporting. We use HubSpot and spreadsheets. Can you start next week, and what will it cost?
At Level 0, the workflow identifies the request, systems mentioned, timing question, and missing details. At Level 1, it drafts a reply that thanks the sender, reflects the reporting problem, asks about report owners and current inputs, and avoids inventing a price or start date.
At Level 2, a person checks the lead record, corrects context, decides whether the opportunity fits, and approves the exact email. The workflow may prepare the message, but the send remains approval-gated in Agentic Workers.
If repeated runs are reliable, the workflow might earn a Level 3 permission to add an internal “reporting inquiry” tag and create a review task. It still should not quote a fee, promise availability, or send the email automatically. Those actions carry different consequences and deserve separate policies.
Failure modes include mistaking a question for purchase intent, using stale CRM details, merging contacts, relying on outdated pricing, or creating duplicate tasks after a retry. Each needs a detection rule and safe fallback.
Run one small, reversible test
Choose 20 recent, non-sensitive inquiries that already have known outcomes. Give the workflow read-only access and ask it to produce drafts in a separate review folder. Do not connect sending.
Have one reviewer score factual accuracy, recipient and context, policy compliance, next step, and editing required. Separate hard failures from style preferences. Then test missing contact data, conflicting notes, and an unsupported pricing request.
Advance only if the evidence meets the criteria you set beforehand. If it does, add one reversible permission, such as creating an internal task for approved categories, and test again with a low volume cap.
One limitation to keep in view
A trust ladder measures behavior on observed cases. It cannot prove how a workflow will handle every new situation. Changes in source data, business policy, connected tools, or customer behavior can make earlier evidence stale. That is why permissions need expiration, periodic review, and automatic return to approval when conditions change.
Credit the source, then own the policy
The draft-first principle was prompted by Corey Ganim’s X post and YouTube walkthrough, published September 15, 2026. He recommends starting communication workflows with human-reviewed drafts, then using real runs and feedback before broadening autonomy. This is source-reported guidance, not proof that a specific workflow is safe.
Your team still owns the permissions, promotion rules, review quality, and consequences. Start with the smallest useful authority, gather evidence, and make every increase in trust explicit.
Audit one repeated workflow and find the safest place to begin.
