ProtoPilot and BioLab Bench: What the Technology Does
On 5 July 2026, MGI Tech’s subsidiary Genoria AI, in collaboration with the Shanghai Artificial Intelligence Laboratory, announced the launch of two interconnected systems: ProtoPilot and BioLab Bench. Together, these tools represent what the developers describe as a “Physical AI” paradigm for the life sciences, a meaningful departure from the text-generation and reasoning tools that have dominated AI adoption in research and commercial laboratory settings over the past several years. The announcement was made via PR Newswire and is supported by technical documentation published through academic preprint channels.
The distinction the developers draw between prior AI tools and this new category is worth taking seriously. Previous AI applications in laboratory contexts largely operated as digital reasoning engines: they could suggest protocols, summarise literature, or generate code in isolation. ProtoPilot is designed to close the loop between that digital reasoning and the physical world of laboratory automation hardware. It translates high-level biological experimental intent, expressed in natural language, into physically executable, verifiable, and reproducible automation code that can be deployed directly on robotic laboratory platforms. This is not a marginal upgrade to existing tools; it represents a structural change in how AI interacts with physical laboratory infrastructure.
For biotechnology firms, pharmaceutical companies, contract research organisations, and the professional services consultancies that support them, the implications are practical and immediate. Autonomous error correction, cross-platform hardware compatibility, and a standardised evaluation framework for AI agents in wet-lab settings change the risk and reliability calculus for laboratory automation investment. Understanding what these systems do, how they perform, and what their limitations are is increasingly relevant for anyone involved in laboratory operations, R&D automation strategy, or technology procurement decisions in the life sciences sector.
ProtoPilot and BioLab Bench: Performance Metrics and Technical Architecture
ProtoPilot’s most directly comparable performance metric is its score of 52.38% on the ProtocolQA benchmark, a public evaluation framework for AI experimental reasoning developed by Future House. Human expert performance on this benchmark sits at approximately 54%. OpenAI’s GPT-5.6 Sol, which represented the strongest competing model tested on the same benchmark, scored 43.5%. The gap between ProtoPilot and the next best AI model is therefore approximately 8.88 percentage points, while the gap between ProtoPilot and human expert performance is less than 2 percentage points. These are statistically meaningful differences in a benchmark specifically constructed to test the kind of reasoning required to design and execute biological experiments.
The system operates on a four-stage execution loop that the developers describe as Design2Protocol, Protocol2Code, Device Execution, and Wet-Lab Feedback. The first stage converts high-level experimental intent into a structured protocol. The second translates that protocol into SDK-compliant code for specific automation hardware. The third stage executes that code on the physical device. The fourth, and arguably most significant, stage feeds the physical outcomes of the wet-lab back into the system, enabling the AI to diagnose failures and autonomously regenerate corrected protocols. This closed-loop architecture is what differentiates ProtoPilot from static code generators: if an antibiotic resistance screening step fails, for example, the system does not require a human operator to identify the cause and rewrite the protocol manually. The AI agent performs that diagnosis and correction autonomously.
Hardware integration performance is one of the more concrete metrics available. On the OpenTrons platform, ProtoPilot achieved an 88.2% code gate pass rate. The comparable figure for OpenTrons-AI, the platform’s own native AI tooling, was 32.4%. That is a difference of 55.8 percentage points on the same hardware, which is a practically significant finding for any organisation weighing up the operational reliability of AI-generated automation code on high-value robotic platforms. The system has also been validated for use with Hamilton STAR and Tecan EVO automation hardware, indicating cross-platform compatibility rather than optimisation for a single vendor ecosystem.
The BioLab Bench evaluation framework was developed alongside ProtoPilot and is described as the industry’s first real-task evaluation framework for AI agents in wet-lab settings. It spans 294 synthetic and molecular biology tasks derived from 98 gold-standard protocols. Crucially, BioLab Bench evaluates agents on three distinct criteria: device-level validity gates, wet-lab expert rubrics, and real physical tests. This is a meaningful methodological departure from text-matching evaluation, which can reward plausible-sounding outputs that fail when executed on physical hardware. The hardware-native data flywheel built into the system means that physical feedback from executed experiments directly trains and updates the AI’s skill library over time, creating a self-improving loop intended to reduce experimental error rates as the system accumulates operational experience.

Australian context: implications for local biotechnology, pharmaceutical, and laboratory consulting sectors
Australia’s biotechnology and pharmaceutical manufacturing sectors have been expanding steadily, supported by federal investment in sovereign manufacturing capability following supply chain vulnerabilities exposed during the pandemic period. The Australian Government’s Medical Research Future Fund and associated manufacturing initiatives have directed substantial capital toward domestic R&D infrastructure, including laboratory automation. For Australian organisations operating or
References and related sources
- Primary source: www.prnewswire.com
- nai500.com
- arxiv.org
- modelscope.cn
How iEnvi can help
iEnvi integrates technology and data-driven approaches into environmental consulting. We monitor AI and technology developments that affect how environmental professionals deliver services to clients.
This is an iEnvi Machete news summary. Prepared by iEnvi to summarise the source article for environmental professionals tracking AI, data, and technology developments that affect consulting and project delivery.
Published: 06 Jul 2026
Need advice on this topic? Speak to an iEnvi expert at info@ienvi.com.au or 1300 043 684, or contact us online.
Need advice on this issue? iEnvi provides practical, senior-led environmental consulting across contaminated land, remediation, ecology and environmental risk.
Contaminated land advice Remediation services Discuss your site Talk to iEnvi