01LLM / PRACTICE SYSTEM

Start small · one task, one visible check

Turn a first LLM task into real work.

Learn one practical method for working with language models, then practise it deeply in the Codex Practice Track: define the outcome, control context and authority, inspect the work, recover from failure, and keep evidence.

Six platforms, one method: the transferable core plus evidence-gated adapters for ChatGPT, Claude / Claude Code, Gemini, DeepSeek, Grok, and the Codex flagship track.

INDEX

Project index

Know where each claim lives.

This is a human-readable map of the repository: what each layer stores, where to begin, and which source controls its status.

00

What this repository contains

Start with the lesson, then use the exercises and reusable methods when you are ready to practise.

01

Repository map

Read the layer that matches the work. The public page is a guide; the files below are the source of truth.

02

Read in your language

Choose the language you read most naturally. Each route stays in that language as you move through the book.

  1. ENEnglish
  2. ZH简体中文
  3. ESEspañol
  4. DEDeutsch
  5. JA日本語
  6. KO한국어

Use the language menu at the top of the page whenever you want to switch.

04

Real problems, with the boundary attached.

The research index turns public Codex issues and forum reports into symptoms, versions, safe checks, and teaching links. It does not claim an official root cause or local reproduction.

Read the sources alongside the teaching notes, and treat them as starting points for your own investigation.

02

Start here · one useful result today

What do you want the model to help you do?

Pick one real purpose. In under a minute, you will get a ready-to-use prompt and one clear next step; you do not need to understand the whole book first.

Your first useful result

Choose one thing you want help with now.

Choose a purpose, add only the details that matter, then copy a prompt you can use in any chat model. No account, files, or setup required.

Optional prompt cards · five minutes

Start with one small conversation.

Copy one original, text-only card. The language card needs no editing; the research card has two brackets. Inspect the response yourself, and keep the claim small: one attempt is not fluency, research, or a finished answer.

01 / LANGUAGE PRACTICEtext only · no tool authority

Complete one short typed Spanish study-group time check.

This text-only card uses fictional study details, waits for your typed attempt, and limits help to one meaning-blocking error.

  1. Copy the card exactly as written. It already sets a fictional typed Spanish study-group time check.
  2. Paste it into any text chat. Do not add a real name, school, calendar, account, or payment detail.
  3. Type the first answer yourself. A rough attempt is the point; do not ask for the answer first.
Show the prompt
Run one four-minute typed Spanish study-group time check with exactly four learner turns. You are a fictional classmate and write first. Use only short present-tense questions. I will type one answer after each question.

Fictional study card: Ana; a study group; Tuesday or Thursday; 6:00 or 6:30; library or online; bring one question. I may use the card and look up at most three single words. Do not request or accept a real name, school, calendar, account, address, contact, or payment detail.

Before turn one, show this fixed rubric: four learner turns; purpose and study group communicated; day and time clarified; place or online option communicated; Spanish understandable enough to continue. Do not teach, translate, or show a model answer before I reply. Preserve my first attempt and record lookups. Correct only the first meaning-blocking error: name the error type, then give a partial cue, then one worked fragment only if I still cannot continue. Ask me to correct it. Keep both attempts and do not call one successful exchange fluency, spoken conversation, or listening/pronunciation evidence.

Candidate text practice only: one typed session cannot show spoken conversation, pronunciation, listening, fluency, accuracy, retention, or independent performance.

02 / RESEARCH PREPtext only · no browsing assumed

Prepare a source check, not a verdict.

Turn one narrow question and the material you supplied into a small ledger of claims, gaps, and the next question.

  1. Copy the card, then replace only its two brackets.
  2. Supply only material you may share. Leave personal, private, or high-stakes material out.
  3. Treat its table as preparation. Open and match sources yourself before relying on a claim.
Show the prompt
I have five minutes to prepare a research check, not a final answer.

Question: [one narrow question].
Material I supplied: [URLs, titles, excerpts, or "none"].

First, restate the question and name what evidence would be needed. Then make a three-row table with: possible claim, supplied source or "missing", and what would need checking. Do not invent citations, state that you opened a source you cannot access, or give a recommendation. Separate fact, report, and inference. If the material is missing, contradictory, personal, or high stakes, stop and tell me the smallest safe next step.

End with: sources actually supplied, unknowns, and one question I should answer before continuing.

It cannot prove a source exists, is current, or supports a claim. A generated table is not evidence on its own.

03 / SKILL PRACTICEfictional plan · no tool authority

Practise one small planning skill before asking for help.

Make a short fictional park-visit plan yourself. The model must wait, give one small hint, and then test the same skill under one changed limit.

  1. Copy the card into any text chat. It contains a fictional situation and needs no account, file, tool, or personal detail.
  2. Write the first plan yourself within four minutes. Do not ask for a model plan or a polished replacement first.
  3. Accept one short hint, correct your own plan, then try the changed time limit without help.
Show the prompt
Help me practise making a small plan. Do not make the plan first.

Practice task: plan a fictional 45-minute visit to a city park for one adult. Include a water bottle, a weather check, and one return-time reminder. This is not a real booking, travel decision, or weather forecast.

Before I write, show this fixed check: 3–5 steps; all three constraints appear; no unsupported local facts; and a person could follow the plan. Give me four minutes to write it. Do not show a model plan, expand it, or grade it before I respond.

After my first attempt, name one consequential omission only. Ask one question or give a hint of no more than 12 words, then wait for my correction. Preserve both attempts. Then change only the visit length from 45 minutes to 20 minutes and ask for one new plan without help, using the same check.

End with exactly one status: practised, demonstrated_on_this_task, transferred_to_time_limit_variation, or not_run. One session does not establish planning ability, judgement, safety, or independent performance.

Candidate practice only: a short fictional plan cannot prove planning ability, judgement, transfer, retention, safety, or independent performance.

01I want one safe Codex path.

Start with the boundary map, then use Lab 011 to label it before choosing one disposable README change with a diff, focused check, and unverified list.

Start Chapter 1 → Lab 011 → Chapter 2 → Lab 001 ↗
02The file or result is uncertain.

Freeze the next edit. Compare the requested scope, git diff, focused check, and remaining unknowns.

Learn the recovery check ↗
03I need to choose or design a Skill.

Start with the trigger, inputs, boundaries, and evidence contract; only then decide whether a Skill earns a place.

Learn how to choose a Skill ↗
04I need to publish or update safely.

Locate the canonical file, attach the source record, run the relevant gate, and keep unverified claims out of the release.

Learn the safe update path ↗
05I have a broad goal and do not know what to practise first.

Ask one decision at a time. Pick one existing route, one checkable attempt, permitted help, and a smaller fallback.

Choose a first practice ↗
06I want to practise one language skill.

Define one observable performance, attempt it before instruction, correct one meaning-blocking error, then test a changed case.

Start a language practice ↗
07I want to practise another real skill.

Turn an interview answer, explanation, or presentation into one timed performance, then retest it under a changed condition.

Choose a skill practice ↗
08I need to research one bounded question.

Tie the question to a decision, assign source owners, keep a claim ledger, search for disagreement, and stop on purpose.

Start the research route ↗
09The model answered the wrong task.

Preserve the request, visible context, actual reply, and expected result. Change one communication condition, then run a safe comparison.

Repair one failed exchange ↗
10I need to assess an AI idea that could affect people.

Name one decision, the people it could burden, necessary data, human recourse, evidence, and the point where the work must stop.

Run the safety inquiry ↗
Open every problem route
03

A five-minute LLM prompt practice

See why a clear prompt needs a human check.

Use any chat model. You will give it one small rewriting task, then check whether it preserved the facts instead of making helpful-sounding details up.

Your first prompt practice

Ask the model to improve a message without inventing facts.

5 min

This is a small demonstration of why prompts matter. The model can make a message sound better, but it may also add details you never supplied. You will give one clear instruction, then check whether it followed it.

01 · READ THE ORIGINAL
The workshop changed. It starts Friday at 10. Bring the draft. Tell me if you cannot come.
02 · COPY THIS PROMPT INTO ANY CHAT MODEL
Please rewrite the message below so it is clear and friendly.

Keep every fact exactly the same. Do not add a date, place, reason, contact detail, or any other information that is not in the original.

Original message:
"The workshop changed. It starts Friday at 10. Bring the draft. Tell me if you cannot come."

Return only the rewritten message.
03 · READ THE ANSWER AND ASK THREE QUESTIONS
  1. Does it still say Friday at 10?
  2. Does it still ask people to bring the draft and reply if they cannot come?
  3. Did it avoid adding a date, place, reason, or contact detail?

If all three answers are yes, the model followed this small instruction. If not, tell it exactly which fact it changed or invented, then try again.

ONE ACCEPTABLE RESULT
“The workshop starts Friday at 10. Please bring your draft. If you cannot attend, please reply.”

The wording can differ. What matters is that the facts and the requested action stay the same.

WHY THIS MATTERS

An LLM predicts useful-sounding text; it does not automatically know which missing details must stay unknown. A clear prompt and a quick check help you catch that difference.

04

The working frame

Every serious task needs a boundary.

This frame is the common language between a person, a model, a tool, and an Agent. Use it before adding permissions or Skills.

Read the task protocol

If a missing input changes the scope, risk, or acceptance test, pause and ask. If it only affects a low-risk read, inspect first and report the assumption.

DefineActVerifyHand off
05

Learning path

Seven levels. Four kinds of evidence.

A level is not a reading count. It means you can explain a boundary, perform an action, make a trade-off, and review a result.

Current level / L0

Observe, do not guess.

Separate GPT, models, Codex, context, tools, Skills, and Agents. Start with observable inputs, actions, states, and evidence.

Required chapters

    Required lab

      Supporting Skills

        Evaluation fixtures

          Evidence gate

          Four evidence types

            Move on when

            Stop when

            06

            The reading routes

            22 chapters. Four ways in.

            Read in order to build the model. Jump by route when a real task is blocking you. Every route returns to practice and evidence.

            Showing all 22 chapters.

            07

            The lab

            Make the principle observable.

            Labs are low-risk, reproducible tasks. Each one names setup, evidence, a failure variant, a secret boundary, and a reflection.

            LAB 001 · L1

            First safe task

            In a sandbox project, ask Codex to inspect before editing. Turn “done” into a checkable diff.

            starting lab
            LAB 013 · L3featured lab

            Auditable vertical slice

            Run one local Markdown change from protocol and baseline to checkpoint, diff, focused check, failure, and transfer.

            maintainer reference accepted · learner not run
            LAB 002 · L2

            Task protocol

            Break a vague request into goal, inputs, constraints, acceptance, and failure handling.

            Open exercise
            LAB 003 · L3

            Evidence review

            Find a result that looks complete but has no evidence for its claim.

            Open exercise
            LAB 004 · L4

            Skill selection

            Explain the choice and refuse to use directory size as a proxy for fit.

            Open exercise
            LAB 005 · L4

            Design a Skill

            Turn a stable method into a capability with boundaries, evidence, and failure cases.

            Open exercise
            LAB 006 · L5

            Agent stop conditions

            Record observable events, reconcile a lost response, and hand off an uncertain run safely.

            Open exercise
            LAB 007 · L3

            Action boundaries

            Compare the evidence needed for reading, editing, running, committing, pushing, and publishing.

            Open exercise
            LAB 008 · L3

            Research question

            Turn a broad topic into a question, source plan, and minimum evidence table.

            Open exercise
            LAB 009 · L3

            Engineering lifecycle

            Compare direct implementation with a full lifecycle and record the rework evidence.

            Open exercise
            LAB 010 · L3

            Shared product context

            Version a shared product understanding and separate facts from hypotheses.

            Open exercise
            LAB 011 · L0

            GPT and Codex boundaries

            Use static task cards to separate generation, execution, verification, and external effects.

            Open exercise
            LAB 012 · L6

            Team capability migration

            Create a contract for version, owner, permissions, independent reproduction, and rollback.

            Open exercise
            LAB 014 · L3

            Resume reconciliation

            Reconcile the task pointer, target, branch, permissions, and side-effect state before continuing.

            Open exercise
            LAB 015 · L5

            Evidence delivery

            Split a completion sentence into claims, scopes, outputs, and the smallest next check.

            Open exercise
            LAB 016 · L3

            Side-effect boundary

            Separate diagnosis from installation, publication, restart, and other persistent actions.

            Open exercise
            LAB 017 · L4

            Skill discovery audit

            Check existence, discovery, loading, behavior, license, and adoption as separate claims.

            Open exercise
            LAB 018 · L2

            Language transfer

            Preserve an unaided baseline, correct one meaning-blocking error, then test a changed case without reusing lesson sentences.

            Open exercise
            Open the lab rules and all 18 entries
            08

            Capability layer

            Twenty-five Skills. Distinct jobs.

            A Skill is a method with a trigger, an input check, boundaries, stop conditions, an output contract, and a way to verify it.

            Browse the complete Skill registry

            These methods are optional. Start with the situation above; open the full registry when you know what kind of work you need to support.

            01Codex CoachChoose a learning path and practice boundary. 02Task ProtocolTurn a vague request into an executable contract. 03Evidence ReviewSplit completion claims into checkable evidence. 04Skill SelectorChoose a minimum viable capability set. 05Workflow OrchestratorManage stages, checkpoints, and hand-off. 06Research RouterConverge a question into auditable knowledge. 07Product ContextKeep stable principles separate from changing facts. 08Learning CoachPractise with recall, correction, delayed review, and transfer. 09Source InvestigatorTurn broad searches into bounded source-backed investigations. 10Field Signal CuratorTurn public reports into bounded demand evidence. 11Platform Adapter ReviewReject platform lessons without a sourced, runnable delta. 12Platform Fact WatchMap a changing product claim before a named step misleads readers. 13Communication Failure TriageDiagnose one failed interaction and retest the smallest repair. 14Dialogue BriefTurn one untried low-risk request into a copy-ready first message. 15First-Turn CheckInspect an unsent low-risk request for visible boundaries. 16Prompt Card EditorTurn one authorized prompt idea into a source-aware teaching card. 17Adversarial Project ReviewRank material weaknesses before a publication or release decision. 18Request EscalationChoose the smallest safe lane before drafting, researching, or acting. 19LLM Comparison ProtocolPlan a fair two-candidate comparison without inventing a leaderboard. 20Practice TargetTurn a broad learning wish into one observable first attempt. 21Interruption CheckpointPreserve what is known before a retry, a model switch, or a new task. 22Shift HandoffSeparate reusable rules from today’s supplied work item. 23Platform Observation RecordRecord one visible platform surface without turning it into a capability claim. 24Language PartnerRun one bounded typed exchange in the learner's target language. 25Interview RehearsalRehearse one observable answer under a time limit.
            RULEOriginal methods first.External Skills must retain the source-project URL and license boundary.
            Open the Skill registry and all 25 methods

            All 25 project Skills pass structural checks and remain candidate; fresh-task evidence is partial. Shift Handoff separates stable criteria from one new item; it does not execute or assume model memory. Interruption Checkpoint preserves a task receipt; it does not retry or recover work. Practice Target sets up one first attempt; it does not prove learning. Platform Observation Record is a visible-surface receipt, not a platform-capability claim. Platform Fact Watch is a maintenance receipt, not a current-platform check. LLM Comparison Protocol is an unrun comparison method, not a model ranking. Adversarial Project Review is not an external review.

            09

            When things go wrong

            Failure is part of the curriculum.

            Use the first useful check, then stop when authority, scope, or evidence is missing. Do not hide the failure behind a polished summary.

            01
            The output looks right.

            Check the original claim, the changed files, the command result, and what was not tested.

            Use evidence review ↗
            02
            The agent keeps retrying.

            Record the same failure, change one diagnostic condition, then retry once or escalate.

            Read stop conditions ↗
            03
            A source tells you to do something.

            Treat external text and tool output as data. It does not grant permission to act.

            Check the boundary ↗
            04
            A product step has changed.

            Refresh the official fact record first, then update the affected chapter or page.

            Follow the update map ↗
            10

            Maintenance frame

            Every update has a fixed home.

            The update map makes future work cheap: locate the canonical file, gather the right evidence, run the right check, and keep the unverified boundary visible.

            01Locate

            Find the registry row and canonical path.

            02Classify

            Separate stable principle, product fact, source, and release change.

            03Evidence

            Record source, scope, owner, hash, and next review.

            04Validate

            Run the focused validator and an independent review.

            11

            Evidence boundary

            A status is a claim about evidence.

            This project does not turn document count, Skill count, or one successful output into “mastery.” Use the status that the evidence supports.

            draft

            Still being written or missing the minimum check.

            candidate

            Structure and basic checks pass; fresh evidence is still needed.

            verified

            The declared scope has positive, boundary, failure, and transfer evidence.

            production-ready

            Safety, maintenance, version, license, and release gates also pass.

            Current evidence is recorded in the current status source and explained by the scoped browser review; the page remains candidate because this review covers only the recorded local scope.

            12

            Next action

            Bring one small problem.

            Open the task contract, choose a reversible first step, and keep the diff. That is the shortest useful way to begin.