MGI Tech's subsidiary Genoria AI, in partnership with the Shanghai Artificial Intelligence Laboratory, launched ProtoPilot and BioLab Bench on July 5 — a self-evolving multi-agent system and its accompanying evaluation framework designed to translate experimental intent into physically executable actions on automated laboratory platforms. The dual release targets the persistent gap between digital reasoning and wet-lab execution, moving beyond AI tools that generate text without closing the loop to physical hardware. Strategically, it positions Genoria AI — established by MGI in April 2026 as a dedicated AI for Science subsidiary — as a direct competitor in the emerging physical AI space for laboratory automation.
ProtoPilot operates across a full experimental chain: design to protocol, protocol to machine code, device execution, and wet-lab feedback. When a step fails, the system diagnoses the error and autonomously regenerates a corrected protocol — a capability demonstrated when it identified and resolved an antibiotic resistance screening failure during a PCA assembly workflow. BioLab Bench addresses a separate but related problem: the absence of any standardised way to evaluate whether an AI agent can actually operate real automation hardware, not merely produce plausible text. The framework assesses agents across three difficulty levels, from fundamental operations to complex multi-step workflows, and tests cross-device transferability across different automated laboratory platforms.
On ProtocolQA, a public benchmark developed by Future House for evaluating AI experimental reasoning, ProtoPilot scored 52.38%, compared with 43.5% for GPT-5.6-sol and a human expert ceiling of 54%. The result, published as a preprint on arXiv in June 2026, places ProtoPilot within two percentage points of expert-level performance on a task that general-purpose frontier models trail by more than ten points. The gap between ProtoPilot and GPT-5.6-sol is notable precisely because it reflects domain-specific agent scaling rather than raw compute advantage.
The competitive landscape for laboratory AI agents currently centres on general-purpose large language models applied post-hoc to experimental workflows, and on standalone lab automation software that lacks adaptive reasoning. Neither category closes the loop between protocol generation and physical execution. Genoria AI's approach differs by embedding failure-driven self-correction directly into the agent architecture and by grounding evaluation in real hardware execution rather than text quality. The company's hardware-native position — MGI's automation platforms are deployed across more than 3,800 users globally — gives it a data advantage that pure-software competitors cannot easily replicate.