15 September 2026 · Proposed acceptance plan; no measurements or tests have run
1. Answer five questions in order
- Can the art direction survive editable 3D, changing light and interaction?
- Does one player action produce a coherent, remembered consequence?
- Does learning a capability make play meaningfully different?
- Do new players enjoy this place and choose to continue?
- Can an automated candidate improve it without damaging what was accepted?
Each gate buys evidence for the next production expense. Do not create all chapter assets, a live-agent service or a studio dashboard before these answers exist.
2. A bounded engine and production proof
Recommended first trial: Unity with Blender, one room section, one working instrument and a single route fragment. Use the same permitted source art for an Unreal comparison only if a decisive rendering/workflow question remains; time-box that comparison and make a choice. Do not demand feature parity across engines.
Verify installed versions and existing hardware first. Choose and record at least one minimum-target Apple Silicon Mac and one minimum-target Windows machine with exact CPU, GPU, RAM, OS, display resolution and power mode. Founder hardware alone cannot define the market floor. Cloud or borrowed Windows hardware may be needed; no such service is purchased by this plan.
Provisional profiles to make the trial concrete: an M2 Mac with 16 GB unified memory, and a Windows PC with Ryzen 5 5600, RTX 3060 12 GB and 16 GB RAM, both at 1080p output. These are proposed benchmark targets, not verified minimum requirements or purchase recommendations. Confirm access and OS/toolchain compatibility before ratifying them. If the scene misses the target, explicitly choose optimization, a lower tier or a changed floor; do not silently benchmark only a faster machine.
First test only 60–90 seconds: look, reach, adjust, see the rain respond, record the observation, save and reload. Proposed camera/input mode is restrained first-person, a short walking area, no head bob, keyboard/mouse and controller. Evaluate control pleasure and comfort before building the full ten-minute route.
Proposed initial budgets
These are starting product targets to ratify after profiling, not achieved results or universal engine limits.
| Area | Proposed target and method |
|---|---|
| Standard native play | Aim for 60 FPS at 1080p output on the named target tier; warm-run p95 frame time at or below 16.7 ms, with p99 and worst spikes reported separately |
| Minimum-quality tier | If 60 is impractical, explicitly support stable 30 FPS at a documented resolution/settings tier; p95 at or below 33.3 ms, preserving the clue and controls |
| Interaction | Immediate visual acknowledgment, target under 100 ms; separately measure input-to-first-visible-response and full action duration |
| Memory | Establish after the first build; retain at least 20% headroom against the agreed per-device process budget; inspect sustained memory growth |
| Save recovery | Exact semantic state before/after normal reload; recover previous valid save from an interrupted write |
| Visual identity | Same approved object, rig and material families in every tested state; zero unreviewed hero-asset substitutions |
| Accessibility | Complete the same discovery by keyboard/controller, without audio, with reduced motion and with readable captions |
Use a fixed ten-minute route, plus cold start and a repeated interaction loop. Measure the actual packaged build, not only the Editor. Include rain, wet reflections, mist, transparent cloth/foliage, characters and captions at the same time. Record thermal/power conditions; alternate candidate order to avoid confounding heat with code quality.
Report GPU and CPU time, resolution/upscaler configuration, p50/p95/p99 frame times, longest stutter, memory high-water mark, startup/load times and capture evidence. A fast headless run cannot prove visual performance. A photoreal offline render cannot prove the native look.
3. Gameplay and save matrix
Required meaningful tests:
- Stillness, Craft, Connection, mixed and uncommitted/baseline completion.
- Wrong prediction, requested hint, assisted timing and revision without a dead end.
- Inspection and questions leave the morning unchanged; committed practice consumes it once.
- Duplicate input, double-click and interrupted action cannot create duplicate progression.
- Slot A and slot B remain independent before and after export/import.
- Imported data is schema/version/type/size validated; prose cannot grant effects or execute code.
- Notebook and companion history agree with recorded observations; testimony remains attributed.
- Saves contain rules/content/schema IDs and stable object IDs. Missing old content produces a supported recovery choice.
- Interrupted save/install/migration; full disk; malformed file; old valid save; rollback to last compatible runtime.
- Leaving and returning causes no undeclared decay or background advance.
Use unit/property tests for rules and actual input-driven play for the player-visible journey. Capture and test only affected scenes on a narrow change, plus the minimum release smoke set; expand testing when changed dependencies or failures justify it.
4. Art acceptance
Review the twelve scenes in the artifact kit at full frame and intended display size. Compare baseline/candidate anonymously where practical. Review a complete ordinary-play recording before selecting highlight stills.
Assess composition, object/character identity, material response, physical contact, readable interaction, comfort, sound coherence, pacing and emotional fit. Mark each as pass, specific issue or untested. Preserve categorical failures separately; no average can cancel a missing hand, unreadable clue or lost save.
Image differences identify where to look. They need tolerance for GPU variance and controlled weather/randomness. Perceptual/model scores can reject suspicious output for inspection but cannot make the final creative decision. An evaluator-created screenshot is required; the implementation agent cannot submit a different prettier build.
5. Human play: small, honest samples
Start with six unfamiliar target players as a diagnostic group, consistent with the existing source plan. Vary game familiarity. Observe first; ask questions after allowing an unprompted opportunity to continue, change approach or finish.
Questions:
- What changed because of what you did?
- What could you do after practicing that you could not do before?
- Which moment stayed with you?
- Where did you want to act but could not?
- Did any part feel slow, confusing, tiring or intrusive?
Proposed initial gate: at least four of six explain a capability's consequence and complete the key interaction without facilitator rescue; no participant loses a save; record exact enjoyment/continuation choices and reasons. This is a revision aid, not statistical evidence of market demand. Silence or lack of complaint is not enjoyment.
If players like the picture but do not understand or care about the action, revise that interaction before expanding the art library or studio automation. If a focused revision still fails, reconsider the encounter itself. More assets, models or tests are not evidence of a better game.
For a candidate-vs-baseline test, randomize order or use separate groups and include the same content exposure. Puzzle knowledge creates carryover, so do not interpret faster second play as a better design. Use both a preference judgment and the causal explanation. Preserve dissent and participant context; return to testing if the candidate helps one group and harms another.
Collect only agreed gameplay evidence and opt-in observations. Do not infer emotional state from private voice/video or retain unrelated conversations. Train no model on participant data without separate, clear permission.
6. Shareability test, after the game earns it
Offer an optional preview from an actual saved discovery. Do not ask everyone to share and then call compliance organic interest. Record the denominator at each step: eligible plays, preview opens, voluntary shares, recipient opens, recipient plays, recipients who want continuation.
A postcard can test attraction; a clip can test comprehension; a playable checkpoint can test whether the invitation delivers. Test these separately where feasible. A private preview has a clear publication choice. No invented activity from friends and no generated scene that overstates available gameplay.
7. Automated-studio rehearsal
Use a real, bounded defect such as an interaction target blocked by its own camera collider. Give the implementer only the permitted files and the accepted baseline. Separately inject a known regression in a disposable candidate and verify the independent evaluator detects it.
Rehearse duplicate job triggers, an expired lock, a failed generator, a partial build, budget exhaustion, evaluator outage, stale base revision, unauthorized path change, unsigned asset and incompatible save. Each stops promotion predictably and retains a useful evidence record.
Complete one repair, one rejected candidate and one restore drill manually before automating coordination. Then run the system in candidate-only mode. Expand its release authority only through a separately approved, versioned policy and repeated evidence; a success count by itself is insufficient.
8. Definition of the first completed commission
A person can enter, act, learn a rain interval, make a discovery, save, close and return on verified native Mac and Windows packages. The ordinary minutes satisfy the art direction. The guide accurately recalls the recorded action. Wrong attempts and accessibility options work. A review receipt includes exact packages, content/asset versions, captures, measured performance, player evidence and remaining limits.
Until then, use precise labels: design draft, generated concept, imported asset, Editor prototype, packaged proof, player-tested candidate or released build.