📊 Full opportunity report: From Concept To Reality: Inkling And The Future Of AI Innovation on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Thinking Machines released Inkling, a massive 975-billion-parameter multimodal AI model, on Hugging Face. Its open availability introduces new possibilities but requires substantial computing resources, as detailed in the original analysis. Key benchmarks and licensing details remain unconfirmed.
Thinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. For more details, see the original analysis. The model is designed to process text, images, and audio within a one-million-token context window, offering new capabilities for complex reasoning tasks. This aligns with the insights from the original analysis. This release is notable because it provides open access to such a large-scale model, although the hardware requirements for running it are extremely high, limiting direct deployment for most users.
Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple modalities, including text, images, audio, and video. Its architecture employs 256 experts, selecting six routed experts per input, and combines global and sliding-window attention techniques. The model’s training and context-length figures have not been independently verified, and its capabilities for video processing are unconfirmed, as native video performance has not been evaluated.
Hugging Face reports that Inkling is available through various inference engines, including Transformers and llama.cpp, with support for BF16 and NVFP4 checkpoints. The BF16 version requires approximately 2 TB of VRAM, while the NVFP4 version needs around 600 GB, making full deployment impractical for typical consumer hardware. Developers can access the model via hosted inference services, but licensing details, safety evaluations, and benchmark results are not yet publicly available.
Implications of Inkling’s Open Multimodal Scale
The release of Inkling signifies a major advancement in large-scale multimodal AI research and application. Its open availability, despite hardware barriers, could accelerate development in fields like scientific research, media analysis, and enterprise data processing that require integrated text, image, and audio reasoning. However, the high hardware demands and lack of detailed benchmarks or licensing clarity mean widespread adoption and evaluation are still limited, shaping the future landscape of multimodal AI deployment.
high performance AI inference hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Large Multimodal Models and Open Releases
Prior to Inkling, most large multimodal models remained closed or required specialized hardware for deployment. The trend toward open models has gained momentum with releases like Meta’s Llama and OpenAI’s GPT variants, but few have approached the scale of Inkling’s 975 billion parameters. The model’s architecture, training data, and capabilities are comparable to other recent efforts in the field, but its open access marks a notable shift towards democratizing large AI models for research and development.
Previous multimodal models, such as OpenAI’s CLIP and DALL·E, demonstrated the potential for integrated visual and language reasoning but with smaller parameter counts and more restricted access. Inkling’s release on Hugging Face aims to bridge this gap, although the hardware requirements remain a significant barrier for most users.
“This model is huge.”
— Hugging Face
large scale GPU for AI training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unverified Aspects of Inkling’s Performance and Licensing
Details about Inkling’s benchmark performance, safety evaluations, and domain-specific fine-tuning capabilities remain undisclosed. The model’s actual efficacy across different modalities, especially video, has not been independently tested or verified. Additionally, licensing terms, usage restrictions, and the availability of training data or source code are not yet clarified, leaving questions about the model’s openness and commercial viability.

Building Live Voice Agents: Deploying Real-Time Multimodal AI Systems to Production
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Expected Developer Testing and Benchmarking Efforts
Developers and organizations with access to the necessary hardware are expected to begin testing Inkling through supported inference platforms to evaluate latency, accuracy, and resource consumption. Independent researchers will likely conduct benchmark comparisons and safety assessments, while further disclosures from Thinking Machines and Hugging Face regarding licensing and safety are anticipated. The results of these evaluations will influence future adoption, fine-tuning, and potential commercialization of the model.
AI model deployment servers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Inkling?
Inkling is a large-scale, 975-billion-parameter multimodal AI model developed by Thinking Machines, capable of processing text, images, and audio within a unified framework, available on Hugging Face.
Can Inkling process video?
While the architecture supports image inputs with a temporal dimension, native video processing performance has not been evaluated or confirmed, so its video capabilities remain unverified.
Is Inkling suitable for running on personal computers?
No. The hardware requirements—up to 2 terabytes of VRAM for BF16—are beyond typical consumer systems, making full deployment impractical without specialized hardware or hosted inference services.
What are the licensing terms for Inkling?
The announcement does not specify licensing details, restrictions, or whether the training data and source code will be publicly available.
What are the next steps for Inkling’s evaluation?
Developers and researchers will conduct performance tests, safety assessments, and benchmark comparisons, with further disclosures expected from the developers regarding licensing and safety standards.
Source: ThorstenMeyerAI.com