The non-thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It gives faster direct answers across text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, and long-context work.
Added Jul 15, 2026
Model weightsContext Window
1.0M
Max Output
32.8K
Avg output tokens (7d)
346 tokens
Input Price (Auto)
$1.00/1M
Output Price (Auto)
$4.05/1M
Cache Read (Auto)
$0.17/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
32.2
Coding Index
52.1
Agentic Index
24.4
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
87.2%
Better than 83% of models compared
HLE
Humanity's Last Exam
31.9%
Better than 83% of models compared
AA-LCR
Long context reasoning evaluation
77.3%
Better than 82% of models compared
GDPval-AA
Economically valuable tasks
33.4%
CritPt
Research-level physics reasoning
5.4%
Coding
SciCode
Python programming for scientific computing
47.0%
Better than 41% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
41.5%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
67.7%
Last updated Sep 7, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Inkling with similar models from the same provider or model family.
Inkling Thinking
thinkingmachines/inkling:thinkingThe thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It reasons natively over text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, long-context work, and controllable thinking effort.
Inkling Small
thinkingmachines/Inkling-SmallThe direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Inkling Small Thinking
thinkingmachines/Inkling-Small:thinkingThe reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
DeepSeek V4 Flash Vision Exp Uncensored
deepseek/deepseek-v4-flash-vision-exp-uncensoredAn uncensored variant of the experimental vision-enabled DeepSeek V4 Flash model for chat, image understanding, reasoning, coding, and tool use, with a 524K context window.
TheDrummer/Artemis v1.1
TheDrummer/Artemis-v1.1TheDrummer's Artemis v1.1 is a Gemma 4 31B fine-tune for creative writing, expressive dialogue, and roleplay, with optional thinking and a 262K context window.
Synth 2.5 Flash Preview
synth-2.5-flashSynth 2.5 Flash Preview is a low-cost text model designed for role-play, character dialogue, and collaborative storytelling.