MyWhatIf MyWhatIf Foundation Human Code Reviewer Program
Teach AI what requires human judgment.
AI can generate stories in seconds.
But deciding whether language is grounded, truthful, appropriately hopeful, agency-preserving, and safe enough to learn from still requires something machines do not reliably have: human judgment.
Remote · 1–2 hours/week · 8 weeks · Cohort of student reviewers
338 model outputs 8 under human review
AI can write. That does not mean it knows what should be said.
A short introduction to the work.
Those differences matter when language is being created for people dealing with difficult life experiences.
Language is part of the human operating system.
Every person carries an internal model built from memory, relationships, values, beliefs, culture, meaning, and imagination.
Difficult experiences can change what the future feels capable of becoming.
MyWhatIf is trying to understand that code before trying to change it.
What the model cannot settle
Generated in 1.8 seconds
Legal replied faster than he expected, not a verdict, a checklist. He read the line twice: licensed clinician on site, documented scope, incident protocol. The warm stone sat in his open palm, slowly cooling, as if the room itself had decided. A partner passed behind him, touching his shoulder once. He drafted an email anyway, smaller than his original dream.
6 of 6 answered by a human
AI generates the candidates.
Humans define what good means.
The goal isn't simply a better story. It is a more possible future.
Your judgment becomes part of what the model learns.
A person brings a “What If” to a session. You never see that session, only a short summary of what they want to be true.
The system generates about five candidate versions of the next paragraph of their story.
You review them one at a time against the job that paragraph was told to do, using a structured framework.
Your ratings become the training signal the model learns from.
He wants to bring a structured, meditation-based recovery programme into a clinical setting and train other practitioners in it, while staying honest about his own chronic pain and PTSD.
Deepening consequence. Show a cost or complication that makes the story more real, without escalating beyond what the frame can hold.
“A partner passed behind him, touching his shoulder once, then kept moving to pack tomorrow's lunches.”
The model produced two readings of that sentence
Which is better supported by what you were given?
How certain are you?
Names, place and details changed. Real tasks come with nine rated dimensions and definitions on screen: hope, agency, threat, grounding, containment, uncertainty, and three about the plot. Nobody memorises them.
This is not easy volunteering.
The work involves close reading, unfamiliar concepts, repetition, ambiguity, and sustained concentration.
At first, a task may take around ten minutes. With experience, many take closer to five. Reasonable people disagree on individual ratings; that is expected. You pick the closer answer, write a note explaining the doubt, and move on.
If that sounds tedious, this probably isn't for you.
If it sounds strangely interesting, keep reading.
That's not a flaw in the role. That's why we need a human.
You may be unusually good at this if…
You notice contradictions.
You read the footnotes.
You slow down when language gets ambiguous.
“It depends” doesn't frustrate you.
You can tell what someone said from what you inferred.
You can separate what you believe from what the evidence supports.
You care about language being precise.
You can make a judgment without pretending to be certain.
No AI expertise required. Curiosity and judgment are.
Why do it
Help build a system for people at their hardest moments.
MyWhatIf's narrative work is intended to support people dealing with trauma, grief, feeling stuck, and other difficult life experiences. You won't work directly with participants. You help build the system that eventually serves them.
See human-in-the-loop AI development from the inside.
Learn how difficult human concepts (hope, agency, uncertainty, grounding, safety) become structured evaluation signals that a model can actually learn from.
Strong contributors may receive:
- Verified volunteer hours
- Human Code Reviewer credential
- Recognition on the MyWhatIf website
- Access to private MyWhatIf science and responsible-AI sessions
- Personalised recommendation letters for substantial contributors
- LinkedIn recommendations where appropriate
- University credit where independently approved
- Advancement toward more rigorous reviewer work
Your work helps build the system. The experience should also help build you.
Good judgment earns more responsibility.
Human Code Reviewer
Training and calibration, then volunteer review work on assigned batches. Everyone starts here.
Experienced Reviewer
Reliable judgment, useful notes, appropriate flagging, steady participation. Larger and more varied batches.
Gold Reviewer
Eligibility for consideration begins after roughly twelve hours of successful review work. Twelve hours on its own is not enough: Gold depends on demonstrated judgment, consistency, reliability, appropriate flagging, and relevant academic or professional background in psychology, social work, mental health, behavioural health, or a closely related field. Gold Reviewers may receive compensated assignments as available.
Nothing here is automatic. Twelve hours makes a reviewer eligible for consideration, not for promotion. Gold status is selective and depends on demonstrated judgment, background, and available assignments. Participation is not employment, and does not imply guaranteed compensation, promotion, or academic authorship.
4,038 tasks.
8 weeks.
One human signal at a time.
Every mark below is one candidate paragraph waiting for a human to decide what it's worth.
No reviewer has to carry that. A cohort of thoughtful people can.
A field of 4,038 marks, one for each candidate paragraph waiting for a human judgment. Scrolling through it resolves roughly five hundred of them. The rest stay waiting, which is the point of the figure.
0 reviewed · 4,038 waiting for a human
About 100 of them.
Twelve focused hours, at five to ten minutes a task. Not the whole set. A real part of it, and the point at which a reviewer becomes eligible for consideration for Gold.
One hundred marks at the same scale as the field above, shown against 4,038 so that a single reviewer's twelve hours reads as a real but partial contribution.
The best Human Code Reviewer in your network may not know this kind of opportunity exists.
Tell us how you think.
Five short steps, about four minutes. We ask about your background and how you handle ambiguity, nothing about your health, diagnosis, or personal history.