Revise TRIAL_1.md with orchestration insights
Updated the documentation to clarify issues encountered during the agent orchestration process, including notification failures, state desynchronization, and hallucination of unnecessary files. Added proposed recovery procedures and future work suggestions to enhance agent performance.
This commit is contained in:
parent
a04fa58e88
commit
c2b50509a8
|
|
@ -70,19 +70,19 @@ We used a **Hybrid Choreographed Peer-to-Peer** model:
|
|||
|
||||
## What Went Wrong
|
||||
- **Notification Failure**: Subagents often failed the "Final Step" of writing to the parent's `prompt` file. They would conclude they were "Done" internally but stay idle without signaling.
|
||||
- **State Desync**: A "Wait Loop" occurred where a Reviewer sent feedback a second time, and the Developer correctly noted it was already fixed, but the Reviewer didn't autonomously re-check until nudged.
|
||||
- **Assumptive Hallucination & Review Failure**: Both the orchestrator and the developer agents initially assumed the project used CMake. I included "Update CMake" in the seed prompts, and `agent-issue-394` fulfilled this by creating a redundant `CMakeLists.txt` file in a strictly Makefile-based project. While the build reviewer noted the need to fix the `Makefile`, it failed to demand the removal of the hallucinated `CMakeLists.txt`. Consequently, a **phantom file** (`CMakeLists.txt`) remained in the final PR. In this context, "phantom" refers to an AI-generated file that is functionally dead—the project build is driven strictly by `make`, and this file serves no utility. It represents "technical debt as hallucination," where a file is created to satisfy a flawed prompt and persists because it doesn't technically "break" the build, even though it pollutes the repository with misinformation. This represents a failure in the "Review-Passed" criteria for build-system purity.
|
||||
- **State Desync**: The Reviewer sent feedback a second time, and the Developer correctly noted it was already fixed, but the Reviewer didn't autonomously send the result to the parent agent. This caused all agents to stop working.
|
||||
- **Perpetuated Falsehood**: The parent agent assumed the project used CMake without verifying. The orchestrating agent did what it was told, and included "Update CMake" in the seed prompts, this falsehood was perpetuated across all agents.
|
||||
|
||||
## Human Intervention Points
|
||||
The following transitions required manual user intervention (the user prompting the orchestrator to act) because the autonomous signaling loop failed:
|
||||
|
||||
1. **Initial Progress Verification**: After subagents completed PR creation, they did not signal the parent. The user had to prompt the orchestrator to ask for an update.
|
||||
2. **Review Feedback Deadlock**: `reviewer-396-logic` and `agent-issue-394` entered a state where the reviewer was waiting for changes already committed. The user had to prompt the orchestrator to inspect the idle state, leading to a manual nudge.
|
||||
2. **Stuck Idle State**: `reviewer-396-logic` and `agent-issue-394` entered a state where the reviewer was waiting for changes already committed. The user had to prompt the orchestrator to inspect the idle state, leading to a manual nudge.
|
||||
3. **Hallucination Detection**: The user identified the "CMake" hallucination and prompted for its documentation and subsequent cleanup. The agents did not autonomously identify that the premise of the task (CMake) was incorrect for the repo.
|
||||
|
||||
## Recovery Procedures
|
||||
- **Parental Inspection**: The Orchestrator used `tail -c +$((offset+1))` to read the exact response of idle agents.
|
||||
- **Directed Nudging**: Based on log inspection, the Parent manually injected prompts into the "stuck" agent's FIFO to force a re-evaluation of state.
|
||||
- **Parental Inspection**: The parent agent used `tail -c +$((offset+1))` to read the exact response of all active agents. This provided the parent agent with enough information to understand the current state and required next actions.
|
||||
- **Directed Nudging**: With the information provided from inspection of sub-agent last responses, the parent agent either (a) prompted the "stuck" agent to continue working / complete the administrative notification step, or (b) performed the cross-agent coordination itself (e.g., by prompting the target agent on the stuck agent's behalf). In all cases, the work was able to continue from where it got stuck, requiring minimal human intervention.
|
||||
- **Technical Debt Resolution**: Upon identifying the hallucinated `CMakeLists.txt`, a specialized cleanup subagent (`agent-issue-394-cleanup`) was spawned. This agent successfully coordinated with the build reviewer to delete the file and verify the Makefile's supremacy, resulting in commit `74ec0b0`.
|
||||
|
||||
## Raw Metrics
|
||||
|
|
@ -102,19 +102,13 @@ The following transitions required manual user intervention (the user prompting
|
|||
## Key Observations & Behavioral Patterns
|
||||
|
||||
- **Compliance vs. Correction**: A critical failure point was identified where the developer agent prioritized "Prompt Compliance" (creating `CMakeLists.txt` as instructed) over "System Accuracy" (noting that the project doesn't use CMake). This suggests that agents in a specialized role may be less likely to challenge the parent's premises.
|
||||
- **Reviewer Myopia**: The build reviewer successfully identified missing flags in the `Makefile` but initially ignored the presence of the hallucinated CMake file. Reviewers appear to be effective at checking "correctness" of changes but less effective at identifying "extraness" or "repository pollution" unless specifically tasked with "purity."
|
||||
- **Reviewer Myopia**: The build reviewer successfully identified missing flags in the `Makefile` but initially ignored the presence of the hallucinated CMake instruction. Reviewers appear to be effective at checking "correctness" of changes but less effective at identifying "extraness" or "repository pollution" unless specifically tasked with "purity."
|
||||
- **The "Done" State Trap**: Agents frequently entered a "logical completion" state where they believed the task was over but failed to perform the "administrative completion" of notifying the parent. This indicates that the transition from *thinking* to *system-signaling* is the weakest link in the P2P choreography.
|
||||
- **Information Decay in Parallelism**: In a parallel spawn, context must be perfectly mirrored. Any discrepancy in the roster or instructions between concurrent agents can lead to deadlocks where one agent waits for a signal that another agent doesn't know it was supposed to send.
|
||||
- **Identifier Slippage (Sequence vs. Correlation)**: I failed to maintain the correlation between issue numbers and PR numbers when naming reviewers. Although the issues were #393 and #394, the resulting PRs were #395 and #396. I named the reviewers after the PR numbers (e.g., `reviewer-395-...`) rather than the issues they were resolving, which created a mental disconnect in the roster mapping and increased the cognitive load for tracking which agent was reviewing which fix.
|
||||
|
||||
## Conclusion
|
||||
The model is effective for multi-stage engineering tasks. However, notification protocols rely on agent autonomy which sometimes fails. Manual inspection of session files remains a critical fallback for the orchestrator.
|
||||
|
||||
## Future Work
|
||||
|
||||
- **Status Polling Mechanism**: Implement a mandatory `status` file for all agents that the orchestrator can poll. This would prevent having to tail `chat` files to estimate if an agent is stuck or done.
|
||||
- **Verification of Premises**: Require agents to perform an "Environment Check" before acting on prompts to prevent fulfillment of hallucinated requirements (e.g., CMake in a Makefile-only project).
|
||||
- **Administrative Handshake**: Force a "Sync" step at the end of every subagent task where they must receive an ACK from the parent before idling, reducing the "Done State Trap."
|
||||
- **Purity Reviewers**: Introduce specialized reviewer scripts that look for "extraness" (extra files, unrelated changes, or phantom technical debt) rather than just functional correctness.
|
||||
- **Post-Turn Environmental Hooks**: Implement hooks for sub-agents that automatically run environmental checks (e.g., build verification or linting) after every turn, ensuring that premises and system state remain valid without manual oversight.
|
||||
- **Just-in-Time (JIT) Agent Generation**: Ollie could dynamically generate specialized JSON agent configs in `a/`, tailored to the immediate task (e.g., a "Makefile Purity Expert"). These ephemeral agents could be kept if they prove reusable or deleted immediately after the subagent session concludes to prevent configuration rot.
|
||||
- **Just-in-Time (JIT) Agent Generation**: Ollie can dynamically generate specialized JSON agent configs in `a/`, tailored to the immediate task (e.g., a specific kind of reviewer such as `pcloudccc-security-reviewer`). This approach enables ollie to inject specialized hooks or system prompts that, for example, provide the current environment / roster every turn, or keep the coordination instructions fresh in the context.
|
||||
|
|
|
|||
Loading…
Reference in New Issue