sby @shipgtm
Gate outbound drafts behind a Jev evaluation check
Drafts each message with Claude, saves it to Gmail, scores it with Jev, and revises or sends it once it passes.
Outcome
Every drafted message scored against explicit pass/fail questions before it is allowed to send.
A revised draft for anything that failed, up to the allowed number of attempts.
Only passed or person-approved messages sent from Gmail; everything else flagged with its reasons.
How it works
- 1
Write the draft
AnthropicDraft the outbound message for the lead.
- 2
Save it as a draft
GoogleSave the message to the sending account's Gmail as an unsent draft.
- 3
Score it
TypeSafe AIAsk yes-or-no questions on outbound quality, factual support (no claim the lead's record doesn't back), and whether the next action is appropriate. Keep each answer with its confidence and, for a no, the reason.
- 4
Revise on failure
AnthropicFor a draft with any answer below
min_confidenceor a no, rewrite it with the reasons in view, update the Gmail draft from step 2, then send it back to step 3. Stop aftermax_revisionsattempts. - 5
Route what still fails
Your agentFlag every draft still failing after
max_revisionsfor a person to fix or approve, with its reasons. - 6
Send what passed
GoogleSend every message that passed step 3, or was approved by a person.
You'll be asked for
How many times a failing draft may be rewritten before it goes to a person
e.g. 2
How sure Jev must be before a pass or fail is used without a person
e.g. 0.8
The file your agent runs
outbound-draft-evaluation-gate.md
Gate outbound drafts behind a Jev evaluation check
Drafts each message with Claude, saves it to Gmail, scores it with Jev, and revises or sends it once it passes.
Set up the tools below, then run the steps in order for the user, carrying each step's results into the next. The run is done when the user has the outcome below.
Outcome
- Every drafted message scored against explicit pass/fail questions before it is allowed to send.
- A revised draft for anything that failed, up to the allowed number of attempts.
- Only passed or person-approved messages sent from Gmail; everything else flagged with its reasons.
Inputs
Ask the user for these before you start.
max_revisions: how many times a failing draft may be rewritten before it goes to a person, e.g. 2min_confidence: how sure Jev must be before a pass or fail is used without a person, e.g. 0.8
Set up
Create a message (Anthropic, tool:anthropic/create-message)
Use the first option your agent supports.
Note: The API call needs the anthropic-version: 2023-06-01 header alongside the key, and max_tokens is required in the body.
CLI (official)
Install the command and sign in with it, then confirm it runs.
curl -fsSL https://claude.ai/install.sh | bash
claude --version
Run claude -p.
API (official)
- Base URL: https://api.anthropic.com
- Endpoint:
POST /v1/messages - Auth: send the header
x-api-key: $ANTHROPIC_API_KEY - Get a key: https://platform.claude.com/settings/keys
Google (tool:google/create-draft, tool:google/send-message)
For each call, use the first option your agent supports that lists it.
Notes:
- Send an email: Needs the
gmail.sendOAuth scope; the message is a base64url-encoded RFC 2822 MIME message, not exposed by the MCP server's tools.
MCP (official, remote)
Add this server to your agent's MCP settings, then sign in when asked.
{ "mcpServers": { "google": { "url": "https://gmailmcp.googleapis.com/mcp/v1" } } }
- Create a draft: call the MCP tool
create_draft
Note: A Developer Preview, open only to Google Workspace accounts in the Workspace Developer Preview Program.
API (official)
- Base URL: https://gmail.googleapis.com
- Send an email:
POST /gmail/v1/users/{userId}/messages/send - Auth: an OAuth access token, sent as
Authorization: Bearer <token>
Answer typed questions (TypeSafe AI, tool:typesafe/answer-typed-questions)
Use the API.
- Base URL: https://api.typesafe.ai
- Endpoint:
POST /v1/systemone - Auth: send the header
Authorization: Bearer $TYPESAFE_API_KEY - Get a key: https://console.typesafe.ai
Note: Put every question about one record in one request. GET /v1/models lists the models and is the cheapest check of a key; a request over the rate limit gets a 429 with a Retry-After header.
Note: Send state, model (jev-latest) and named questions, each with type, instructions and criteria: a choice maps labels to descriptions, a score lists its levels in order, a noul's is optional. Read a score by its most likely level; send unsure answers to a person.
Before step 1, confirm access to each service with its cheapest read-only call, like a list or a search. Never send or change anything to test access.
Steps
- Write the draft with Create a message (Anthropic). Draft the outbound message for the lead.
- Save it as a draft with Create a draft (Google). Save the message to the sending account's Gmail as an unsent draft.
- Score it with Answer typed questions (TypeSafe AI). Ask yes-or-no questions on outbound quality, factual support (no claim the lead's record doesn't back), and whether the next action is appropriate. Keep each answer with its confidence and, for a no, the reason.
- Revise on failure with Create a message (Anthropic). For a draft with any answer below
min_confidenceor a no, rewrite it with the reasons in view, update the Gmail draft from step 2, then send it back to step 3. Stop aftermax_revisionsattempts. - Route what still fails yourself. Flag every draft still failing after
max_revisionsfor a person to fix or approve, with its reasons. - Send what passed with Send an email (Google). Send every message that passed step 3, or was approved by a person.
Notes
Keep each pass/fail question narrow enough for a careful editor to judge in a second; a vague question like "sounds on-brand" produces an unreliable score.
Adapted from ShipGTM's GTM agent evaluation guide, which pairs Jev's evaluation model with Claude for drafting and Gmail for sending and replies.
Rules
- Use only the services set up above. The read-only calls they need, like listing ids or polling for results, are fine.
- Ask the user before anything that sends messages, costs money, or changes data, and say how many records it touches. One approval covers a batch the user has seen.
- Never print API keys.