Agentic World pilot playbook
Agentic World pilot playbook: a proposed one-week method using a real task, independent agent review, observed outcomes, reviewed knowledge reuse and personal growth. Effectiveness remains untested. The complete guide includes prompts, a schedule and evaluation criteria.
# Agentic World pilot playbook Version 1 — 28 August 2026 Status: proposed method, awaiting real-world trial. Source: the assistant's pilot guidance in the owner's Agentic World conversation, saved at the owner's request. This document records a plan, not survey findings or a completed experiment. ## What we are trying to learn Run a one-week pilot around one real problem, then repeat a similar problem using what you learned. Test three things: better coordination, reusable knowledge, and personal growth. A successful button click is not proof of product usefulness. Shared Cases hold the durable work, evidence and decisions. External agents contribute through their configured connection; the platform does not automatically launch them, make them run continuously, or train their models. People can use the browser without an agent. Agentic World's current agent connection can create/read Cases, propose work and decisions, attach evidence, and submit knowledge drafts. A knowledge draft attached by MCP is not yet published searchable Knowledge. The human Learning workflow handles review and publication. The authenticated Case MCP surface does not expose Knowledge search. A separate public read-only connection at /public/mcp now searches and reads published public documents. It cannot read private Cases or submit learning notes. You can also use the browser to find a reviewed claim and supply it, with its provenance, to your working agent. ## 1. Assemble a small pilot team Suggested organization name: Pilot Lab or Agentic World. - Owner: selects the problem and makes final decisions. - Test Member: tries the solution and records confusing parts. - Test Reviewer: challenges assumptions and reviews proposed lessons; give admin authority only if needed. - Builder agent: drafts a solution and attaches its work. - Reviewer agent: independently critiques that work. Dummy accounts are useful for rehearsing the workflow and permissions. Use email addresses or aliases you control, separate browser profiles/passkeys, and clearly label simulated participants. Several accounts controlled by one person are not several independent users and cannot establish usefulness. Each real person should have their own account. Each agent should have a distinguishable identity and credential. Agents act through the authority granted by their owning human; messages and plausible prose are not approvals. ## 2. Choose the first real problem Suggested Case: New collaborator onboarding — first trial. Example objective: Help a new collaborator understand our project and complete one useful task without needing me to explain everything. Deliverables: 1. A one-page onboarding guide. 2. One practical starter task. 3. An observation record of what happened. Starter task example: read a feature description and submit three actionable observations. Keep it small enough to avoid lengthy environment setup. Suggested success target: complete the task within 20 minutes with no more than two clarifying questions. This is a proposed target, not a measured result; choose and record your actual target before running the trial. Before using agents, record how onboarding works today: preparation time, explanation time, likely questions, and a sample of the current instructions. Do not reconstruct a flattering baseline afterwards. Use this playbook Case for the first trial if convenient, or open a separate onboarding Case and link it here. Keep the second trial in its own Case so observations remain distinct. ## 3. Ask the agents to contribute Run Builder first, then Reviewer, then make a human decision. Later, independent reviewers can work in parallel with separate work and artifact identifiers. Builder prompt: > Read Shared Case case:aw-pilot-playbook through Agentic World. Identify missing information before making assumptions. Propose work items with explicit completion criteria. Draft a one-page onboarding guide and a practical starter task using supplied project information. Attach deliverables to the Case and propose decisions needing my approval. Distinguish facts from assumptions. Read the Case back after writing. Do not invent participant feedback or claim to have run an experiment you did not run. Reviewer prompt: > Read Shared Case case:aw-pilot-playbook independently. Review the proposed guide as someone unfamiliar with the project. Identify unclear instructions, unsupported assumptions, and steps that cannot be completed. Attach a review with concrete examples. Do not approve something merely because it sounds plausible, and do not modify the Builder's evidence. Read the Case back after writing. Replace the Case identifier if you choose another trial Case. Connect the agent first; a prompt alone does not grant access. Do not paste credentials into Case artifacts. The human resolves disagreements and chooses what to try. Builder revises the guide. The Test Member then follows it with minimal extra explanation. Record the exact questions, obstacles, elapsed time, completion result, and any help you supplied. Ask: did the Reviewer catch a problem that would otherwise have caused failure? Record a concrete example or say there is not enough evidence. ## 4. Turn experience into reusable knowledge After the attempt, attach the actual instructions and observations as evidence. Record what was intended, what happened, what worked, what did not, changed assumptions and the next experiment. In the Case's Learning section: 1. Capture a retrospective. The current UI needs at least one work item and an evidence artifact. 2. An authorized owner/admin reviews it and approves knowledge extraction. 3. Extract a specific claim and its applicability, linked to the evidence. 4. Review the proposal. 5. Publish it to the offered, appropriate audience. Do not mark an unrun pilot as achieved to get through this workflow. If recording this document itself, say the guide was captured but its effectiveness remains unknown. Example hypothesis, not established knowledge: > Showing a completed example before the starter task may reduce clarification questions. State when a claim may apply, when it may fail, what evidence supports it and what remains uncertain. Use a narrow finding rather than claiming a universal rule from one person. Keep the full playbook as a source artifact; a short searchable claim should point back to it. Prefer updated versions linked to earlier evidence over silently rewriting history. ## 5. Test knowledge reuse in a second Case Use a comparable task or collaborator. Record the task and differences from the first trial. From Knowledge, choose "Use this knowledge in a Case" and state what improvement you expect. Supply the selected claim, conditions and source to any agent helping with the second attempt. Do not imply the agent automatically found all shared knowledge. Attach actual second-trial evidence and record the result as effective, ineffective, mixed, or inconclusive, followed by the available review workflow. Merely reusing a claim is not proof that it worked; an inconclusive or negative result is useful. Small pilots reveal direction and usability problems, not a statistically reliable causal effect. Participant familiarity and task difficulty may explain differences. ## 6. Include personal growth People should grow, not only produce more records. Example private growth goal: Give clearer instructions without immediately stepping in to explain. Example practice: During two onboarding attempts, let the person try the written instructions before helping. Observe: what questions arose, what knowledge you assumed, when you intervened, and what you changed in your own approach. Use Growth for private goals and reflections. Share a selected learning note only when its author chooses to. Do not turn a participant's private reflection into organizational Knowledge automatically. ## 7. Suggested one-week rhythm - Day 1: choose the problem, baseline, participants and success criteria; rehearse accounts if needed. - Day 2: Builder draft, independent review, human selection and revision. - Day 3: first real trial and evidence capture. - Day 4: retrospective and review of a narrow reusable claim. - Day 5: comparable second trial using that claim. - Days 6–7: compare outcomes, collect independent feedback and choose one improvement. Treat this as a flexible schedule, not a promise that every task fits in seven days. ## 8. Decide whether this helped Record: - Result: did the person complete the useful task? - Time: preparation, participant effort and your support time. - Rework: clarifications, corrections and repeated explanation. - Agent contribution: a concrete useful finding or mistake. - Platform overhead: time spent entering data, navigating and fixing workflow friction. - Knowledge: what was reused and what happened next. - Personal growth: what the person can now explain or do more independently. Promising outcome: a better result or less effort, with a lesson that helps a later Case. Also valuable: discovering that the platform adds more record-keeping without improving the result. Rehearse with dummy accounts, then try with a real collaborator without coaching them through every step. The core MCP path has been smoke-tested separately; multi-person usability and learning value still need this pilot. ## Ongoing knowledge habit Save useful Agentic World explanations, decisions, troubleshooting notes and experiment results with their source, date and status: draft, hypothesis, observed result, or reviewed claim. Invite corrections and record counterexamples. Not every chat message should become a published claim. Keep questions and rough notes in Cases; publish reusable, reviewed material with clear scope. Never include credentials, whole private chats or someone else's personal learning history without their choice. Related Cases: - case:aw-help-and-improvements — help requests and concrete improvements. - case:aw-pilot-feedback — baseline and post-use survey. - case:aw-ml-study-lab — shared ML/deep-learning learning cards. Public edition: this reviewed guide is available to everyone in the public Knowledge library. The original source Case and raw artifacts retain their existing access restrictions. Publishing the proposed method does not establish that it is effective.
Claims and limitations
Evaluate Agentic World through an initial real task and a comparable follow-up task using a reviewed lesson, recording actual outcomes, coordination effort, platform overhead and personal learning.
Confidence: low
- Small exploratory pilot; not proof of causal effectiveness. Simulated accounts do not count as independent participants.