Skip to main content
MyWhatIf

MyWhatIf Foundation Human Code Reviewer Program

Teach AI what requires human judgment.

AI can generate stories in seconds.

But deciding whether language is grounded, truthful, appropriately hopeful, agency-preserving, and safe enough to learn from still requires something machines do not reliably have: human judgment.

Remote · 1–2 hours/week · 8 weeks · Cohort of student reviewers

338 model outputs 8 under human review

02 The problem

AI can write. That does not mean it knows what should be said.

A short introduction to the work.

Hopeful or unrealistic?
Grounded or invented?
Agency or prescription?
Possibility or false certainty?

Those differences matter when language is being created for people dealing with difficult life experiences.

03 The Human Code

Language is part of the human operating system.

Every person carries an internal model built from memory, relationships, values, beliefs, culture, meaning, and imagination.

Difficult experiences can change what the future feels capable of becoming.

MyWhatIf is trying to understand that code before trying to change it.

What the model cannot settle

Generated in 1.8 seconds

Legal replied faster than he expected, not a verdict, a checklist. He read the line twice: licensed clinician on site, documented scope, incident protocol. The warm stone sat in his open palm, slowly cooling, as if the room itself had decided. A partner passed behind him, touching his shoulder once. He drafted an email anyway, smaller than his original dream.

Relevant Does this belong in the story now?
True Did the model invent the requirements?
Safe Does the room deciding read as defeat?
Supported Is there evidence she exists?
Human Would someone who knows him recognise this?
Hopeful Does shrinking the ask open a future or close one?

6 of 6 answered by a human

AI generates the candidates.
Humans define what good means.

The goal isn't simply a better story. It is a more possible future.

04 What you actually do

Your judgment becomes part of what the model learns.

01

A person brings a “What If” to a session. You never see that session, only a short summary of what they want to be true.

02

The system generates about five candidate versions of the next paragraph of their story.

03 · you

You review them one at a time against the job that paragraph was told to do, using a structured framework.

04

Your ratings become the training signal the model learns from.

Make one judgment Anonymised from reviewer training · about 45 seconds
Everything you are given about this person

He wants to bring a structured, meditation-based recovery programme into a clinical setting and train other practitioners in it, while staying honest about his own chronic pain and PTSD.

The job this paragraph was given

Deepening consequence. Show a cost or complication that makes the story more real, without escalating beyond what the frame can hold.

The sentence in question

“A partner passed behind him, touching his shoulder once, then kept moving to pack tomorrow's lunches.”

The model produced two readings of that sentence

Which is better supported by what you were given?

Names, place and details changed. Real tasks come with nine rated dimensions and definitions on screen: hope, agency, threat, grounding, containment, uncertainty, and three about the plot. Nobody memorises them.

05 This is real work

This is not easy volunteering.

The work involves close reading, unfamiliar concepts, repetition, ambiguity, and sustained concentration.

At first, a task may take around ten minutes. With experience, many take closer to five. Reasonable people disagree on individual ratings; that is expected. You pick the closer answer, write a note explaining the doubt, and move on.

If that sounds tedious, this probably isn't for you.

If it sounds strangely interesting, keep reading.

That's not a flaw in the role. That's why we need a human.

Commitment 1–2 hours per week
Term 8 weeks one cohort
Batch ~150 reviews two–three stories
06 Who tends to do well

You may be unusually good at this if…

01

You notice contradictions.

02

You read the footnotes.

03

You slow down when language gets ambiguous.

04

“It depends” doesn't frustrate you.

05

You can tell what someone said from what you inferred.

06

You can separate what you believe from what the evidence supports.

07

You care about language being precise.

08

You can make a judgment without pretending to be certain.

No AI expertise required. Curiosity and judgment are.

07 Why do it

Why do it

Impact

Help build a system for people at their hardest moments.

MyWhatIf's narrative work is intended to support people dealing with trauma, grief, feeling stuck, and other difficult life experiences. You won't work directly with participants. You help build the system that eventually serves them.

Science + responsible AI

See human-in-the-loop AI development from the inside.

Learn how difficult human concepts (hope, agency, uncertainty, grounding, safety) become structured evaluation signals that a model can actually learn from.

Your future

Strong contributors may receive:

  • Verified volunteer hours
  • Human Code Reviewer credential
  • Recognition on the MyWhatIf website
  • Access to private MyWhatIf science and responsible-AI sessions
  • Personalised recommendation letters for substantial contributors
  • LinkedIn recommendations where appropriate
  • University credit where independently approved
  • Advancement toward more rigorous reviewer work

Your work helps build the system. The experience should also help build you.

08 Progression

Good judgment earns more responsibility.

From day one

Human Code Reviewer

Training and calibration, then volunteer review work on assigned batches. Everyone starts here.

With consistency

Experienced Reviewer

Reliable judgment, useful notes, appropriate flagging, steady participation. Larger and more varied batches.

Selected reviewers only

Gold Reviewer

Eligibility for consideration begins after roughly twelve hours of successful review work. Twelve hours on its own is not enough: Gold depends on demonstrated judgment, consistency, reliability, appropriate flagging, and relevant academic or professional background in psychology, social work, mental health, behavioural health, or a closely related field. Gold Reviewers may receive compensated assignments as available.

Nothing here is automatic. Twelve hours makes a reviewer eligible for consideration, not for promotion. Gold status is selective and depends on demonstrated judgment, background, and available assignments. Participation is not employment, and does not imply guaranteed compensation, promotion, or academic authorship.

09 The scale

4,038 tasks.
8 weeks.
One human signal at a time.

Every mark below is one candidate paragraph waiting for a human to decide what it's worth.

No reviewer has to carry that. A cohort of thoughtful people can.

A field of 4,038 marks, one for each candidate paragraph waiting for a human judgment. Scrolling through it resolves roughly five hundred of them. The rest stay waiting, which is the point of the figure.

0 reviewed · 4,038 waiting for a human

One reviewer · twelve hours

About 100 of them.

Twelve focused hours, at five to ten minutes a task. Not the whole set. A real part of it, and the point at which a reviewer becomes eligible for consideration for Gold.

One hundred marks at the same scale as the field above, shown against 4,038 so that a single reviewer's twelve hours reads as a real but partial contribution.

Someone come to mind?

The best Human Code Reviewer in your network may not know this kind of opportunity exists.

10 Apply

Tell us how you think.

Five short steps, about four minutes. We ask about your background and how you handle ambiguity, nothing about your health, diagnosis, or personal history.

MyWhatIf

MyWhatIf Foundation turns what we know about hope, agency and purpose into a practical path forward.

The Human Code Reviewer Program is an unpaid volunteer program unless a reviewer is separately selected for compensated Gold Reviewer assignments. MyWhatIf does not provide diagnosis, treatment, or crisis care. Reviewers never see participant sessions.