← / → or Space · F fullscreen

Stop prompting.
Start delegating.

Using Claude as a junior employee, not a chat window
A working session, August 2026

Consistent AI output comes from a written standard, not a better prompt

What changed

The work moved from phrasing the ask well to writing down the procedure once, in a file.

The model stopped being the variable. Your briefing became the variable.

What it means for your week

You stop re-explaining the same background every morning.

You review a decision instead of rewriting a draft's voice.

Two staff members asking the same thing get the same answer.

Think grant reports, donor thank-yous, board memos, and program updates.

What stays yours

Judgment, the final sign-off, and the disclosure.

The draft arrives done. The decision is still a person's.

Nothing here signs a grant report for you.

The governing thought

"Consistent, on-brand AI output does not come from better prompts; it comes from writing your standard down in a file the model reads every time, which is what turns a chatbot into a junior employee who drafts while you decide."

2
Executive summary
01
Section 01

The chatbot habit is a 2023 habit

Why your instinct to open a chat window caps what you get

3

The craft moved from writing clever prompts to writing down how the work gets done

Prompt engineering

Phrase the ask well. Output changes when the phrasing changes, so nothing is repeatable.

2022–2023
2024–2025

Context engineering

Curate what the model sees. The material you give it matters more than the wording you use.

Skills and harness

Write the procedure down once, in a file, and reuse it. The harness is what the AI can reach.

2025–2026
2026

Loop engineering

Define a goal and a check, then let it iterate until it passes the check.

Emerging practice

Emerging: graph engineering

The next step past loops: many loops wired into a dependency map, so work routes itself from one to the next without a person moving it.

Source: Anthropic Engineering, practitioner terminology, 2026 [S004]
These are practitioner terms, not formal industry standards.
4
Terminology

A chatbot answers the question you typed; a junior employee delivers the thing you asked for

From

How it goes now

  • You open a blank chat window and paste the same background again
  • What comes back is nearly right and sounds like nobody who works here
  • You rewrite the voice by hand, every time
  • Tomorrow you start over from nothing
  • Two staff ask the same question and get two different answers
The difference is not the model. It is the briefing.
To

How it goes when it is briefed

  • The standard lives in a file, not in somebody's head
  • It drafts to that standard on the first pass
  • You review the decision, not the sentences
  • The second person to ask gets the same output as the first
  • What it learned about your voice outlasts whoever set it up
5
From / to
02
Section 02

What you are actually working with

The three layers underneath, the three doors in, and how a prompt becomes an agent

6

Three layers decide your output: the model thinks, the harness acts, and your skill sets the standard

The stack read top to bottom, as it acts on your request
SkillYours to control

The written procedure for one recurring job at your organisation. It reads your brand files to apply the standard: your voice, your rules, your approved examples.

One job, written downReads your brand filesApplies the standard
HarnessWhat it can see and touch

Your files and folders, plus Google Drive, Microsoft 365, and Slack, plus the ability to run several steps and check its own work before handing anything back.

Your foldersGoogle DriveMicrosoft 365 & SlackMulti-step + self-check
ModelNot yours to tune

The raw reasoning ability. The same for everyone who pays for it; a better prompt does not change it.

Raw reasoningSame for every userNot prompt-tunable
The skill layer is the only one that is yours. It is also the one that decides whether the output is consistent.
Source: Anthropic Engineering, harness design, 2026 [S005][S003]
7
The three layers

Chat, Cowork, and Code are three different jobs, and Cowork is the one built for your team

What it isWhat it can touchWho it is forWhat you get backWhere it breaks
Claude
(chat app)
A conversation window Only what you paste into it Anyone, thinking out loud An answer you copy elsewhere Nothing persists, so you re-brief it every time
Claude
Cowork
Claude doing the work in your files, on desktop or web Local files and folders, plus Microsoft 365, Google Drive, Slack and Amplitude1 Staff who do recurring work, no coding required A finished draft in the right place, with every step it took shown to you; can run on a daily, weekly or monthly schedule It needs the standard written down first, or it guesses
Claude Code An agent that runs in a terminal A code repository Engineers Working software The wrong door for programme staff
Most nonprofit work sits in the middle row. Most nonprofit staff are still in the top row.
1. Connector availability depends on your plan and admin settings.
Source: Claude Cowork product page, accessed August 2026 [S003]
8
Three doors in

Prompts become skills, skills chain into an agent, and each step removes work you were repeating

1

Prompt

A one-off ask, typed fresh. Useful, and it varies every time.

2

Skill

That same ask written down as a procedure in a file, with your voice rules and one example of good. Now the output repeats.

3

Chained skills

Intake, then draft, then check against the rules, then package for review. Each step is a small skill; the chain is the workflow.

4

Local agent

A trigger plus a goal runs the whole chain without you starting it.

The same job, at each rung
You paste the transcript and ask for a thank-you note.
donor-thankyou.md runs it: your voice, your rules, one good example.
Transcript in, draft out, checked against the rules, filed for review.
The recording lands and the draft is waiting when you sit down.
An agent is not a smarter model. It is your own skills with a trigger attached.
9
From prompt to agent
03
Section 03

Where consistency comes from, and where it stops

The engine, the refiner, and the three hard edges

10

Context, memory, and brand files solve different problems, and only brand files survive the conversation

Markdown is a text file with headings in it. That is the entire technology underneath everything on this slide. If you can write a Word document, you can write the thing that makes your AI output consistent.

Context

Everything inside the conversation you are having right now: what you pasted, what it already said, the file you dragged in. Powerful, and it disappears the moment you close the window.

Memory

What the tool carries between sessions on its own. Convenient, and you did not write it. You cannot easily see it, edit it, or hand it to a colleague.

Brand files

Plain markdown you own, in a folder: voice, boilerplate, approved examples, who signs off on what, read fresh by every person on staff. This is the standard.

Only the third one gives two different staff members the same answer on a Tuesday.
voice.md
# Voice
Warm, specific, never corporate. Short sentences.
## Always
- Name the actual program, site, or person
- Say what the money or the hour actually bought
- Write to one reader, not to "our supporters"
## Never
- "Your generous contribution"
- "We are thrilled to announce"
- Claim an outcome we have not measured
## Sign-off
Program staff sign donor notes. The ED signs anything about money.
That is the whole technology. A text file, in a folder, that every draft reads first.
Where consistency comes from
11

Give it a trigger and a goal, and the loop turns a rough first draft into an output you would have signed

Fails the check, loop back
Trigger
The call recording lands in the folder. Nobody starts anything.
Goal and standard
What done means, plus the brand files it must write to.
First draft
Fast, rough, wrong in a few places. The part people mistake for finished.
Check
Every claim cited, voice matches, under 400 words, no unapproved promises.
Revise
Fixes what the check flagged, not a vague sense of better.
Output
Arrives ready for a human decision.
  • The measurable part is the check. 'Make it better' does nothing; 'every claim must cite a transcript line' does.
  • The loop is only ever as good as the check you can write down.
  • This is the long game: you record the call, and the pipeline hands you a draft to decide on.
The checks shown are illustrative. Yours is whatever you can state plainly enough to test.
The loop
12

A local agent is safe to experiment with; one that runs on a server needs its own logins, storage, and a named owner

The risk is not the model going rogue. It is an unattended process holding your logins and touching real records while nobody is watching.

Local. Start here
  • Runs on your machine, under your own login
  • Uses only files you already have access to
  • Every step it takes is shown, so you can check its work
  • Worst case is a bad draft you delete
Production. Different conversation
  • Runs on a server, using its own login instead of yours
  • Needs credentials it can hold and use without you present
  • Needs somewhere to store data, and rules about what it may touch
  • Needs monitoring, and a named person accountable when it touches a donor record

Stay local until the output is boring. Boring is the signal that it is ready to run without you.

The failure you will actually meet

It will invent a number. A plausible participant count or success rate that appears nowhere in your files. Check every figure before it leaves the building.

It will use last year's facts. If the file says the program runs Tuesdays and it moved to Thursdays, it will keep saying Tuesdays. Your files are only as current as you keep them.

It will sound certain either way. Confidence is not a signal of accuracy. The tone is identical when it is right and when it is wrong.

13
Where consistency stops

Claude has started marking its output invisibly, so plan to disclose rather than hope nobody checks

What changed

New Claude models launching on or after 2 August 2026 support machine-readable marking at launch. Existing models are being retrofitted during a transition period.

How it works

An imperceptible watermark is woven into generated text and survives copying and pasting. Generated .svg, .png and .jpg files also carry a signed record showing they came from Claude (an open standard called C2PA).

Where it applies

The API (the developer connection), the Claude app, Claude Code, Claude Cowork and Claude Tag (Claude inside Slack). Worldwide, not only in Europe.

Why

EU AI Act Article 50(2) requires generative AI output to be marked and detectable as artificial. The obligation to mark it sits with Anthropic, not with you; your call is how you disclose when AI helped.

What it does not do

A detected mark says content was processed by Claude. It does not say who wrote it, and it can disappear when text is heavily edited, paraphrased or translated.

Write your disclosure sentence now, while how it reads is still your choice.

Source: Anthropic Help Center, how Claude marks AI-generated content; EU AI Act Article 50(2), accessed August 2026 [S001][S002][S006]
14
Where consistency stops

Brief it in writing, keep the standard in a file, and stay local until the output is boring

1

Brief it once, in writing, like a junior employee

Stop re-explaining yourself every morning to something with no memory of yesterday.

2

Your standard has to live in a file, not in someone's head

Otherwise it walks out the door when they do.

3

Cowork is your door

Chat is for thinking out loud, Code is for engineers, Cowork is for the work you actually repeat.

4

An agent is a trigger plus a goal plus your own skills

It is not a smarter model, and waiting for a smarter model will not get you there.

5

Stay local until the output is boring, and write your disclosure line before anyone asks for it

Both of those are cheap now and expensive later.

None of this is about working faster. It is about walking into the meeting with the draft already done.
15
Where consistency stops