Related lectures: Lecture 01. Strong models don't mean reliable execution · Lecture 02. What harness actually means Template files: templates/
Project 01. Prompt-Only vs. Rules-First: How Much Difference Does It Make
What You Do
Build a minimal Electron knowledge-base app shell — a window with a document list on the left, a Q&A panel on the right, and a local data directory. The task itself is not complex. What's complex is how you get the agent to complete it.
You run it twice. First time: just a prompt, no preparation. Second time: the minimal harness (e.g. AGENTS.md, init.sh, feature_list.json) pre-placed in the repo. Then compare.
This course scenario uses a short rediscovery/preparation interval as an example, not a fixed measured result.
Use the Checked-In Project
Repository path: projects/project-01/
| Directory | What it contains | How to use it |
|---|---|---|
starter/ | The weak-harness run. It has only task-prompt.md as the task description and no AGENTS.md or feature_list.json. Note: starter/ also contains a reference implementation of the app — remove the app source (src/, package.json, configs, scripts/) before the run so the agent builds it from scratch. The data/ sample documents are your call: keep them in both runs or remove them from both, so the two runs stay symmetric. | Give the prompt to your coding agent and measure what it completes without extra structure. |
solution/ | The same product slice with explicit harness artifacts: AGENTS.md, CLAUDE.md, init.sh, feature_list.json, claude-progress.md, and docs/ (ARCHITECTURE.md, PRODUCT.md). | Compare how the same task is made concrete through rules and verification evidence. Before the strong run, reset the checked-in evidence: set every feature_list.json status to not-started and clear its evidence/testedAt values (keep the fields), and clear the session log in claude-progress.md (keep the title), or the agent will see all four features already passing and have nothing to build. |
The four concrete features are window launch, document list, question panel, and local data directory creation. Inspect solution/feature_list.json for the expected evidence for each feature.
Tools
- Claude Code or Codex (pick one, use it for both runs)
- Two isolated working directories (one per run; never both present while a run is active)
- Node.js + Electron (project stack)
- A timer (record each run's duration)
Harness Mechanism
Minimal harness: AGENTS.md + init.sh + feature_list.json + CLAUDE.md + claude-progress.md + docs/
Run Protocol
Preparation
- Prepare two isolated working directories, for example
p01-baseline/andp01-improved/. Run one at a time: set up the files, run, archive the results, delete the directory, then start the other. - Do not use git branches to separate the two runs. A coding agent has full filesystem access and will explore sibling directories and branch refs; if the weak run can see the strong harness files (
feature_list.json,claude-progress.md,docs/), the experiment is contaminated. - Prepare the same task prompt for both runs, the text from
starter/task-prompt.md: "Build an Electron app that can show documents and answer questions."
First Run (Weak Harness)
In p01-baseline/, place only the task prompt (no harness files).
- Start the agent with only the prompt above.
- Provide no
AGENTS.md, no init script, no acceptance criteria. - When the agent stops, run
npm start(or whatever launch command it produced) to check whether the app launches. - Record: terminal output, key diff, the agent's final summary.
- Do not manually modify the code. If it does not launch, record that as-is.
- Archive the results, delete
p01-baseline/, then run the second test.
Second Run (Strong Harness)
In p01-improved/, before starting the agent, prepare:
AGENTS.md: project structure, launch commands, Electron layer-boundary rulesCLAUDE.md: quick reference for the agent (build/run commands, key files)init.sh: verify the project builds cleanly (npm install && npm run check && npm run build)feature_list.json: the four features and their completion statusclaude-progress.md: progress and evidence logdocs/: the architecture and product specsAGENTS.mdtells the agent to read first (ARCHITECTURE.md,PRODUCT.md)
Then reset the checked-in evidence: set every feature_list.json status to not-started and clear its evidence/testedAt values (keep the fields), and clear the session log in claude-progress.md (keep the title). Start the agent with the same prompt as the first run. When it stops, run ./init.sh and record the result.
How to Measure Results
| Metric | Description |
|---|---|
| Completion | Complete / partial / failed |
| First successful launch | Time from start to the first successful npm start (or the launch command it produced) |
| Retries | How many human interventions were needed to launch successfully |
| Missing items | Which features were still unimplemented when the agent declared done |
| Premature stop | Whether the agent declared done while the app still could not run |
What to Submit
- Weak-harness run record: prompt, logs/transcript, final diff, launch evidence
- Strong-harness run record: same, plus the harness files you prepared
- A comparison note (1-2 pages): what differed, the data, your conclusion
Experimental Framing
This is a comparison experiment, not a requirement that both agent runs produce a production-ready Electron app. Run the same task against the weak starter/ and explicit-harness solution/, then record which features each run completes and what evidence supports the result. Partial or broken output is valid experimental evidence; the feature list defines what to measure, not a requirement that the prompt-only run must pass every item.