microlab is an experiment in working collegially with frontier AI on a real, instrumented system.
The attraction isn't autonomous coding. It's being able to hold an idea in conversation with something technically capable, have it criticised and developed, and move directly from that conversation into working machinery — without having to keep every implementation detail in my own head.
I want my ideas to survive criticism and rigour and take on a life of their own.
The relationship is conversational rather than adversarial. The agent frequently changes my mind; I push back where local knowledge, recency or intent says otherwise. Good technical judgement is respected regardless of whether it came from the human or the model.
Sometimes the implementation gets ahead of my immediate understanding. That's acceptable. I recognise good patterns and can drill down until I understand them when I need to. The useful delegation is cognitive load, not intellectual ownership.
Course correction is cheap: nah, that's too complicated; why are we doing this; what broke; try the simpler seam. The conversation continues and the machinery changes with it.
The model/provider is not important to the operating model. Today Orange is Claude and Blue is Codex/OpenAI. They are persistent system identities, not roles that depend on personality.
Conversation carries intent, criticism and judgement. It is not the source of truth.
The lab's authoritative engineering state lives in the local Forge: code, issues, PRs, Decisions, Findings and history. A model, provider or chat can disappear without taking the lab's identity with it.
The desired operating loop is roughly:
idea → conversation → work in the Forge → implementation → evidence → criticism/correction → reconciliation
The human remains able to intervene throughout, but human-in-the-loop does not mean human-in-every-loop. Routine engineering should not turn into an approval queue.
The harness was not adopted as an agent framework. It has been derived from first principles while operating the lab.
A useful test is simple: remove Blue tomorrow. The architecture is still wanted for Orange. Remove Orange and replace it with another frontier model; the architecture should still make sense.
The harness therefore isn't fundamentally a multi-agent architecture. It is a way of putting an AI agent inside a real system without making the agent the system.
The two-agent experiment began mostly from curiosity. It has been useful precisely because it exposed assumptions about identity, attribution and authority that were easy to miss with one shared account.