Business

MGI Tech's ProtoPilot closes the gap between AI reasoning and wet-lab execution

MGI Tech's ProtoPilot closes the gap between AI reasoning and wet-lab execution

MGI Tech's subsidiary Genoria AI, in partnership with the Shanghai Artificial Intelligence Laboratory, launched ProtoPilot and BioLab Bench on July 5 — a self-evolving multi-agent system and its accompanying evaluation framework designed to translate experimental intent into physically executable actions on automated laboratory platforms. The dual release targets the persistent gap between digital reasoning and wet-lab execution, moving beyond AI tools that generate text without closing the loop to physical hardware. Strategically, it positions Genoria AI — established by MGI in April 2026 as a dedicated AI for Science subsidiary — as a direct competitor in the emerging physical AI space for laboratory automation.

ProtoPilot operates across a full experimental chain: design to protocol, protocol to machine code, device execution, and wet-lab feedback. When a step fails, the system diagnoses the error and autonomously regenerates a corrected protocol — a capability demonstrated when it identified and resolved an antibiotic resistance screening failure during a PCA assembly workflow. BioLab Bench addresses a separate but related problem: the absence of any standardised way to evaluate whether an AI agent can actually operate real automation hardware, not merely produce plausible text. The framework assesses agents across three difficulty levels, from fundamental operations to complex multi-step workflows, and tests cross-device transferability across different automated laboratory platforms.

On ProtocolQA, a public benchmark developed by Future House for evaluating AI experimental reasoning, ProtoPilot scored 52.38%, compared with 43.5% for GPT-5.6-sol and a human expert ceiling of 54%. The result, published as a preprint on arXiv in June 2026, places ProtoPilot within two percentage points of expert-level performance on a task that general-purpose frontier models trail by more than ten points. The gap between ProtoPilot and GPT-5.6-sol is notable precisely because it reflects domain-specific agent scaling rather than raw compute advantage.

The competitive landscape for laboratory AI agents currently centres on general-purpose large language models applied post-hoc to experimental workflows, and on standalone lab automation software that lacks adaptive reasoning. Neither category closes the loop between protocol generation and physical execution. Genoria AI's approach differs by embedding failure-driven self-correction directly into the agent architecture and by grounding evaluation in real hardware execution rather than text quality. The company's hardware-native position — MGI's automation platforms are deployed across more than 3,800 users globally — gives it a data advantage that pure-software competitors cannot easily replicate.

The AllSci BriefSystematic R&D and deal news. Daily.

The roadmap centres on building what Genoria AI describes as 7×24 unattended intelligent laboratories, where agents continuously accumulate real experimental data, failure cases, and expert validations to improve through physical rather than text-based training. CEO Dr. Yang Meng has framed this as a deliberate alternative to the compute-scaling strategies pursued by general-purpose AI developers, relying instead on closed-loop data engineering grounded in real device constraints and wet-lab results. Whether the approach can sustain its benchmark advantage as general-purpose models continue to scale remains an open question, but the hardware integration depth and proprietary experimental data corpus represent a structural moat that warrants continued attention from both the AI life sciences and laboratory automation sectors.


This article was generated with AI assistance and reviewed and edited by the AllSci editorial team Explore more at AllSci News: https://allsci.com/news/


Spot something wrong? Report an issue with this article