Thinking Machines releases Inkling-Small, an open-source reasoning model firms can run in-house
Thinking Machines, the AI startup founded by former OpenAI chief technology officer Mira Murati, has released Inkling-Small, a 276-billion-parameter open-source multimodal reasoning model licensed under the permissive Apache 2.0 terms. The release lands just two weeks after the company’s flagship 975-billion-parameter Inkling model, and the smaller sibling retains roughly 97 per cent of that predecessor’s benchmark performance while requiring only about a quarter of the compute and deployment footprint. For environmental consultancies, engineering firms, and law practices that handle sensitive site data, this is not a minor incremental release. It is a demonstration that top-tier reasoning capability no longer requires either a hyperscale cloud budget or a willingness to send confidential client data to a third-party API.
For Australian environmental professionals and their clients, including developers, planning lawyers, and local councils, the practical question this raises is straightforward: can a mid-sized consultancy now run a genuinely capable AI model on infrastructure it owns and controls, without the data sovereignty and confidentiality exposure that comes with sending site history, contamination assessment data, or hydrogeological modelling inputs to an external vendor’s servers. Inkling-Small suggests the answer is increasingly yes.
This matters because the environmental consulting sector routinely deals with information that clients, regulators, and legal counsel expect to remain confidential, including preliminary site investigation findings, groundwater monitoring data tied to active remediation works, and commercially sensitive due diligence reports prepared ahead of property transactions. As open-weight models close the performance gap with closed frontier systems, the barrier to running this kind of analysis in-house, on private hardware, is falling faster than most firms have planned for.
Key details
Inkling-Small carries 276 billion total parameters but uses a sparse mixture-of-experts style architecture that activates only 12 billion parameters per token during inference. That is a substantial reduction from the 41 billion active parameters required by the original 975-billion-parameter Inkling model, and it is this drop in active parameter count, rather than the headline total parameter figure, that drives the roughly fourfold reduction in compute and hardware footprint quoted by Thinking Machines.
On the third-party Artificial Analysis Intelligence Index, an independently run benchmark suite used to compare large language models across reasoning, coding, and knowledge tasks, Inkling-Small scored 40 out of 100, just one point behind the flagship Inkling model’s score of 41. Notably, Inkling-Small outperformed its larger predecessor on several specific coding and mathematical reasoning benchmarks, suggesting the smaller model’s training or architecture is not simply a compressed version of the larger one but has been tuned to compete on tasks that matter most for technical and professional services work.
The model supports native multimodal input, accepting text, images, and audio, with text as the output modality, and features a 1-million-token context window. A context window of that size is large enough to hold lengthy technical reports, extensive site history records, or multiple regulatory documents in a single processing session without truncation, which is a meaningful capability for any workflow involving long-form environmental or legal documentation.
Critically, the model is released in full under the Apache 2.0 licence, a permissive open-source terms of use that allows commercial deployment, modification, and redistribution with minimal restriction. Model weights are available directly from Hugging Face, and Thinking Machines has also made native fine-tuning support available through its Tinker API, allowing organisations to adapt the base model to domain-specific tasks using their own proprietary training data rather than relying solely on generic prompting.

Business and professional services implications for Australian practice
This development is international in origin, but its relevance to Australian professional services, including environmental consulting, engineering, and legal practices, is direct rather than indirect. Australian firms handling contaminated land assessments, planning certificate advice, or environmental due diligence reports frequently operate under contractual and, in some cases, statutory confidentiality obligations regarding client site data. Sending that data to an offshore third-party API for AI-assisted analysis raises genuine questions under Australian privacy law and under client confidentiality agreements, particularly where the data includes personal information, commercially sensitive transaction details, or information subject to legal professional privilege in matters heading toward litigation or expert witness proceedings.
Open-weight models with a permissive licence like Apache 2.0 change that calculus. A firm can, in principle, download Inkling-Small’s weights, deploy the model on private cloud infrastructure hosted within Australia, or on-premises GPU hardware, and process sensitive documentation without that data ever leaving the organisation’s own environment. This is a meaningfully different governance posture from using a hosted commercial API, where prompts and documents are transmitted to and processed by a third party’s servers, often located overseas, subject to that vendor’s own data retention and jurisdictional terms.
For Australian mid-tier consultancies and in-house technical teams that have previously been priced out of running frontier-class AI models due to the multi-million-dollar compute budgets associated with models in the 500-billion to 1-trillion parameter range, the reduced active parameter count changes the economics considerably. A model activating 12 billion parameters per token sits within reach of a modest on-premises GPU cluster or an Australian-hosted private cloud instance, rather than demanding hyperscale infrastructure. Firms weighing a deployment should still budget for the engineering effort required to host, secure, and maintain a model of this size, and should validate its outputs against professional standards before relying on it in client-facing work. But the cost gap between running a capable reasoning model in-house and paying for a hosted commercial API has narrowed sharply, and firms that handle confidential site data now have a credible, permissively licensed option for keeping that data entirely within their own walls.
References and related sources
- Primary source: venturebeat.com
- https://venturebeat.com/ai/thinking-machines-debuts-inkling-small-open-source-ai
- NEPM Assessment of Site Contamination
How iEnvi can help
iEnvi integrates technology and data-driven approaches into environmental consulting. We monitor AI and technology developments that affect how environmental professionals deliver services to clients.
This is an iEnvi Machete news summary. Prepared by iEnvi to summarise the source article for environmental professionals tracking AI, data, and technology developments that affect consulting and project delivery.
Published: 03 Aug 2026
Need advice on this topic? Speak to an iEnvi expert at info@ienvi.com.au or 1300 043 684, or contact us online.
Need advice on this issue? iEnvi provides practical, senior-led environmental consulting across contaminated land, remediation, ecology and environmental risk.
PSI services Contaminated land services Remediation services Groundwater services Talk to iEnvi