MENU

Thinking Machines Lab Unveils Inkling, a New Open-Weight 975B MoE LLM, Achieving Superior Benchmark Performance Over GLM-5.2

Sebastian Raschka Germany
Overview
Thinking Machines Lab has announced Inkling, a new open-weight Large Language Model (LLM) with approximately one trillion parameters. This 975B Mixture-of-Experts (MoE) model demonstrates solid performance in reported benchmarks, notably surpassing GLM-5.2 with 79.8% on IFBench (vs. 73.3%) and 43.9% on SimpleQA Verified (vs. 38.1%). Featuring an architecture with 41B active parameters and a context window up to 1M tokens, Inkling provides a powerful new tool for the large-scale AI research community.
In Depth

Key Findings

Thinking Machines Lab has unveiled ‘Inkling,’ a new open-weight Large Language Model (LLM) boasting approximately one trillion parameters. This model demonstrated superior performance across several benchmarks compared to existing leading models, achieving notably high scores of 79.8% on IFBench (compared to GLM-5.2’s 73.3%) and 43.9% on SimpleQA Verified (compared to GLM-5.2’s 38.1%), indicating significant outperformance over GLM-5.2.

Technical Details

Inkling adopts a sparse Mixture-of-Experts (MoE) architecture with 975 billion parameters. MoE models are characterized by their ability to dynamically activate different ‘expert’ networks based on the input, allowing for a substantial increase in parameter count while keeping inference computational costs manageable.

  • Parameter Count: The total parameter count is a massive 975 billion (approximately 1 trillion), but the number of active parameters activated at any one time is limited to 41 billion. This enables efficient inference while maintaining access to an extensive knowledge base.
  • Context Window: It supports a vast context window of up to 1 million tokens. This means the model can perform exceptionally well in understanding, summarizing very long documents, or working with large codebases.
  • Benchmark Performance:
    • IFBench: Achieved a score of 79.8%, significantly outperforming GLM-5.2’s 73.3%. IFBench evaluates factual consistency, which is crucial for measuring the accuracy and reliability of AI-generated information.
    • SimpleQA Verified: Scored 43.9%, surpassing GLM-5.2’s 38.1%. This indicates strong question-answering capabilities, particularly for accurate knowledge retrieval in response to simple queries.

These results suggest that Inkling is competitive with, or even superior to, leading existing open-weight models in general knowledge, reasoning, and factual question-answering tasks.

Background & Context

In the field of large language models (LLMs), while proprietary models like OpenAI’s GPT series and Google’s Gemini garner significant attention, open-weight models such as Meta’s Llama series and GLM series are also rapidly advancing. Open-weight models are indispensable for fostering transparency and innovation in AI research, as they allow researchers and developers to freely investigate and customize the models’ internal structures. Thinking Machines Lab’s announcement of Inkling demonstrates that open-weight LLMs can rival or even surpass commercial models in specific tasks, injecting new vibrancy into the open AI community. The adoption of the MoE architecture, in particular, is a significant trend for balancing model scale and efficiency.

Strategic Significance & Outlook

The advent of high-performance open-weight LLMs like Inkling will further accelerate the democratization of AI research and development. Developers and companies will be able to integrate advanced AI functionalities into their applications without incurring licensing fees or API usage costs for foundation models. This offers significant opportunities for fine-tuning LLMs for specific industry verticals and pioneering new research areas. Furthermore, its vast context window and high factual consistency will facilitate applications across diverse business use cases, including long-form content generation, advanced information retrieval, and complex problem-solving. The future utilization of Inkling within the open-source community and the new innovations it will generate are keenly anticipated.

Source: https://sebastianraschka.com/blog/2026/inkling-architecture-benchmark-notes.html

Get our weekly technology intelligence — free

Receive an infographic that lets you judge at a glance whether each field’s analysis report is worth reading.

Subscribe Free — Weekly Tech Intelligence

By subscribing, you’ll receive Troy-Technical’s weekly technology intelligence newsletter.

  • Your email and selected fields are used only to deliver the newsletter.
  • We never share your information with third parties.
  • You can unsubscribe anytime via the link in each email.

See our Privacy Policy for details.

Takes about a minute · Unsubscribe anytime

Let's share this post !

Author of this article

Comments

To comment

TOC