Skip to main content

Kolibri 1: Aleph Alpha's Sovereign Open-Weight LLM

Aleph Alpha has released Kolibri 1, a 78B mixture-of-experts model with 3.46B active parameters, Apache 2.0 weights, and a 1M-token context, trained from scratch in Germany.
Oct 5, 2026  · 14 min read

Explore with AI

ChatGPTClaudePerplexity

It's been a while since we've seen a new company enter the LLM race. But on October 3, Aleph Alpha put its new model Kolibri 1 on Hugging Face under an Apache 2.0 license and asked Europe to take sovereign AI seriously.

The pitch is that European governments and companies can run a capable model on their own hardware, even fully offline, instead of depending on a US provider's API.

Kolibri 1 is a mixture-of-experts (MoE) model with 78B total parameters and 3.46B active per token, trained from scratch by Aleph Alpha Research on 768 NVIDIA B200 GPUs in data centers in Germany and Finland.

It handles German and English only, supports a 1,048,576-token context window (Aleph Alpha recommends staying at or below 262,144 tokens for serving efficiency), and ships as generally available rather than a preview. The pre-training run covered 20 trillion tokens, followed by 3.44T in mid-training and 201B for long-context extension.

In this article, I'll cover everything new with Kolibri 1, looking at the key features, digging into the benchmarks, and going through what it costs to actually run.

TL;DR

  • Kolibri 1 is Aleph Alpha's open-weight bilingual model, Apache 2.0 licensed and trained from scratch in Europe with no base-model dependencies.
  • It leads its stated comparison set on math and German reasoning, and trails Qwen3.6 35B-A3B on closed-book knowledge, MMLU-Pro, and multi-turn tool use.
  • With 3.46B active parameters, it serves fast, but the full FP8 weights still need roughly 78 GB of GPU memory.
  • There is no published hosted API price and no confirmed API model ID as of October 5, 2026. You self-host from Hugging Face with Aleph Alpha's vLLM plugin, or you talk to sales.
  • Switch to it if German-language quality, Apache 2.0 licensing, or EU AI Act compliance is a hard requirement. Otherwise, the open-weight competition is still broader.

What Is Kolibri 1?

Kolibri 1 is Aleph Alpha's specialized bilingual large language model (LLM), released on October 3, 2026, under Apache 2.0. It's a single generally available model, with no smaller or larger siblings.

Under the hood, it's a 50-layer mixture-of-experts transformer.

Each layer holds 384 small specialist sub-networks called experts, and a router sends every token to just 6 of them, plus 1 shared expert that always runs. That's how 78.1B total parameters shrink to 3.46B active per token (about 4.4%), so the model is built for cheap serving rather than maximum capacity.

Two more design choices keep it fast and light on memory:

  • Attention: four out of every five layers only look at the nearest 512 tokens (sliding-window attention), while every fifth layer looks at the whole context. This is what makes very long inputs affordable.
  • Precision: the weights are stored in 8-bit floating point (FP8), roughly half the size of a standard 16-bit model. The sensitive parts (embeddings, output head, normalization layers, and the router) stay in higher-precision bfloat16.

The headline result Aleph Alpha is pushing is efficiency rather than raw capability.

Kolibri 1 averages 71% on its German benchmark suite while decoding roughly 47,000 bytes/s of text per GPU, against 68% at about 24,000 bytes/s for Nemotron Super A12B.

The model has a knowledge cutoff of June 18, 2026, for both languages, and it is text-only, with no image, audio, or video input or output.

If you want to see where the German capability actually comes from, the published data mix is unusually transparent for an open-weight release.

Domain Pre-training Mid-training Long-context extension
English web and documents 43.4% 1.3% 15.8%
German web and documents 23.4% 2.1% 2.6%
Code 13.6% 16.3% 13.2%
Agentic code and tool use ~0% 15.4% 5.6%
OCR'd PDFs 0% 0% 33.3%

The table shows only the web, code, agentic, and PDF rows. The model card also lists English and German instruction and reasoning data, STEM documents, QA, and specialized legal and public-sector data. Counting all English and German rows, the model card summarizes the pre-training mix as roughly 62.5% English, 23.9% German, and 13.6% code.

kolibri-1-model-card

Kolibri 1 Key Features

Most of what distinguishes Kolibri 1 comes down to two decisions: train the German side properly instead of translating it, and keep the active parameter count low enough that a single H200 can serve the thing.

German that wasn't bolted on afterwards

Kolibri 1 narrows the gap between German and English more than you'd expect, which is unusual for an open-weight model of this size.

Its overall benchmark average is 75.5 in English and 70.8 in German.

Aleph Alpha trained a 128,000-token vocabulary with a new tokenizer algorithm called UniBPE, which keeps the greedy bottom-up merges of byte-pair encoding but uses the Unigram objective to pick which merge enters the vocabulary.

The stated payoff is German compression of 4.90 bytes/token, against 4.35 for GPT-5's tokenizer, with English holding at 4.58 bytes/token (GPT-5's tokenizer gets 4.67).

That matters for cost as much as quality.

Better compression means fewer tokens for the same German document, so a 50-page Bescheid (an official German administrative notice) costs less to process and leaves more room in context.

German made up 23.4% of the pre-training mix against 43.4% for English web and documents, so the capability is paid for in data, not in a fine-tune bolted on after the main training run.

A 1M-token context window with an honest asterisk

The model advertises 1,048,576 tokens. It handles 262,144 tokens natively after a 201B-token long-context training stage, and stretches to 1M by extrapolation, meaning it was never trained at that length.

Aleph Alpha itself recommends staying at or below 262,144 tokens for serving efficiency. RULER, a test of whether a model can find and use specific details buried in very long inputs, shows the degradation for the base model: 67.9 at 128K and 63.2 at 1M, against a cited Qwen3.5 35B-A3B Base result of 89.9 at 128K.

In fairness, Kolibri's 1M score still beats Nemotron 3 Nano (58.5) and Qwen3.5 (57.5) at that length, so it fades less than they do from a lower starting point.

I'd treat the 1M figure as a ceiling you can technically hit rather than a working budget.

If your retrieval-augmented generation (RAG) pipeline stuffs 400K tokens of scanned PDFs converted to text (OCR) into a prompt and expects the model to reliably find one specific detail, test it before you commit. Notably, 33.3% of the long-context extension mix was OCR'd PDFs, so scanned-document workloads are clearly the target use case.

Controllable reasoning mode and native tool calling

Reasoning mode means the model works through a problem step by step before answering, which improves accuracy but uses more tokens.

In Kolibri 1, it's a switch you control rather than a separate model tier, and tool calling (the model deciding to call an external function or API, such as a database lookup) is native.

Aleph Alpha lists multi-step reasoning, RAG, agentic tool calling, coding, and bilingual assistant work as the intended uses, and the mid-training mix put 15.4% into agentic code and tool use, up from essentially zero during pre-training.

One anecdotal user report, running the model in FP8 on a single RTX Pro 6000 workstation GPU, measured around 170 tokens per second and noted a tendency to spend a lot of tokens deliberating.

Budget for that if you leave reasoning on by default in a latency-sensitive path.

Deployment you own end-to-end

Kolibri 1 was trained from scratch, so it isn't a fine-tune or copy of another company's model.

That matters for licensing and for knowing exactly what went into it.

Combined with Apache 2.0 weights, it means you can run Kolibri 1 air-gapped (fully disconnected from the internet), fine-tune it, and ship it inside a commercial product without asking anyone.

The EU AI Act and its General-Purpose AI (GPAI) Code of Practice set the EU's rules for general-purpose models, including documentation of training data.

Aleph Alpha is a signatory of the Code of Practice and publishes a detailed model card covering training-data sources, decontamination steps (removing benchmark questions from the training data), and a point of contact for rightsholders, which is the documentation an EU AI Act compliance review will ask for.

For public-sector buyers in Germany and the wider EU, that paperwork is often the deciding factor rather than a benchmark delta.

It's also the part no US lab can copy quickly.

Documented limits you should plan around

Aleph Alpha's model card lists the limitations openly, and they are worth reading before a procurement decision:

  • Text only, with no image, audio, or video input or output.
  • Implicit world knowledge stops at June 18, 2026, though tool use can supply newer information.
  • Systemic and political bias is not ruled out, and the model was not tuned to hold a particular viewpoint. The training data also includes some material generated with Chinese language models, which carry known political bias. Aleph Alpha says it worked to reduce this.
  • The model is built for human-AI collaboration. In decision-support systems it belongs on the advisory side, not as the deciding component.

For any high-stakes deployment, Aleph Alpha states that additional guardrails and compliance with Regulation (EU) 2024/1689 are required.

How Does Kolibri 1 Perform on the Benchmarks?

Kolibri 1's benchmark profile is lopsided in a useful way: it leads its stated comparison set on competition math and German reasoning, and loses to Qwen3.6 35B-A3B on closed-book knowledge, MMLU-Pro, and most tool-use suites.

Aleph Alpha compares against Qwen3.6 35B-A3B, Mistral Small 4 119B-A6B, and Nemotron 3 Super 120B-A12B.

The model card also lists Qwen3.8 27B, a dense model (one that uses all its parameters for every token, unlike an MoE) that scores higher across the board (overall German 79.9 against Kolibri's 70.8) at roughly eight times the active parameters.

I've left it out of the tables below, but keep it in mind when reading the efficiency claims.

In those model names, the number after the "A" is the active parameters per token, so 35B-A3B means 35B total and 3B active. Here is what the benchmarks in the tables below measure:

  • AIME: competition-level math problems from a US high-school olympiad qualifier.
  • GPQA Diamond: very hard, expert-written science questions in biology, physics, and chemistry.
  • Humanity's Last Exam (HLE): a deliberately brutal, broad test built from questions experts wrote to stump AI.
  • MMLU-Pro and MMLU-ProX: broad multi-subject knowledge questions. "CoT" means the model reasons step by step first, and ProX is the multilingual version used for German.
  • AA-Omniscience: tests factual recall, and whether the model admits when it doesn't know instead of making something up.
  • Tau2-Bench, Tau3-Bench, and BFCL: simulated customer-service tasks where the model uses tools across a conversation, and a test of how accurately it calls functions.
  • LiveCodeBench, SWE-bench Verified, and Terminal-Bench: writing code from scratch, fixing real GitHub bugs, and working in a command line.

kolibri-1-benchmarks

Reasoning and math

Math is where Kolibri 1 is clearly ahead.

It scores 96% on AIME 2026 (EN) against 91% for Qwen3.6 35B-A3B, 90.4% for Nemotron 3 Super, and 83.1% for Mistral Small 4, and 96.9% on AIME 2025 (EN) against Qwen3.6's 84.6%.

On GPQA Diamond (EN) it takes 84.3% versus 83.4%, and on Humanity's Last Exam (EN) 21.5% versus 21.1%, both close enough to call a tie.

The weak spot in the same table is knowledge.

The AA-Omniscience Index (public set) is scored on a scale where negative numbers are common and higher is better. Kolibri 1 manages -32.8 against -12.5 for Qwen3.6 and -24.0 for Mistral Small 4, though it edges Nemotron 3 Super's -36.5. A separate AA-Omniscience accuracy figure comes in at 14.8%, the lowest of the four models.

Its non-hallucination rate on the same benchmark is a more flattering 44.0%, ahead of Nemotron (13.9%) and Mistral (34.7%) but behind Qwen3.6 (56.7%), which fits Aleph Alpha's abstention training.

A 3.46B active-parameter model has limited room to memorize facts, so pair it with retrieval rather than trusting closed-book recall.

Benchmark (EN) Kolibri 1 Qwen3.6 35B-A3B Mistral Small 4 119B-A6B Nemotron 3 Super 120B-A12B
AIME 2026 96% 91% 83.1% 90.4%
AIME 2025 96.9% 84.6% 79.8% 91.7%
GPQA Diamond 84.3% 83.4% 74.7% 78%
Humanity's Last Exam 21.5% 21.1% 9.7% 20.6%
MMLU-Pro CoT 80% 84.3% 80.4% 82.7%
AA-Omniscience Index (higher is better) -32.8 -12.5 -24.0 -36.5

German-language tasks

Kolibri 1 takes the German math and science benchmarks and loses the German knowledge ones, which is the same shape as its English profile.

It scores 90% on AIME 2026 (DE) against 84.4% for Qwen3.6, and 81.3% on GPQA Diamond (DE) against 80.6%. On MMLU-ProX CoT (DE) it drops to 75.5% while Qwen3.6 reaches 81.9% and Nemotron 3 Super 79.7%, and on Humanity's Last Exam (DE) it manages 15.9% against Nemotron's 22.3%.

What I find genuinely interesting is the small English-to-German gap.

Kolibri loses 6 points moving AIME 2026 from English to German and 3 points on GPQA Diamond, which is a narrower drop than you'd expect from a model that treats German as a translation target.

Benchmark (DE) Kolibri 1 Qwen3.6 35B-A3B Mistral Small 4 119B-A6B Nemotron 3 Super 120B-A12B
AIME 2026 90% 84.4% 78.5% 87.5%
AIME 2025 87.5% 82.9% 72.3% 85.6%
GPQA Diamond 81.3% 80.6% 72.9% 76.6%
MMLU-ProX CoT 75.5% 81.9% 70.7% 79.7%
Humanity's Last Exam 15.9% 20.5% 10.5% 22.3%

Tool calling and agentic work

This is the softest part of the release.

On BFCL v4 overall, Kolibri 1 scores 61.4% against 67.2% for Qwen3.6, and on Tau2-Bench (Telecom) it hits 94.7% against Qwen3.6's 99.1%. It does edge Qwen3.6 on Tau2-Bench (Airline), 76.7% to 70.7%.

A separately reported BFCL v3 multi-turn figure of 39.8 points (53.5 for Qwen3.6) shows the same weakness: single tool calls are fine, long conversations with repeated tool round-trips are not.

The outlier runs the other way.

On Tau3-Bench (Banking), Kolibri 1 scores 38.1% while Qwen3.6 gets 10.6%, Nemotron 3 Super 15.5%, and Mistral Small 4 just 5.7%. That is roughly a 2.5x margin over the next-best model in this comparison set (3.6x over Qwen3.6) on a finance-flavored agent benchmark, and it is the single result in this release I would most want independently reproduced.

Benchmark Kolibri 1 Qwen3.6 35B-A3B Mistral Small 4 119B-A6B Nemotron 3 Super 120B-A12B
Tau2-Bench (Telecom) 94.7% 99.1% 41.5% 68.1%
Tau2-Bench (Retail) 69.9% 71.6% 62.9% 67.5%
Tau2-Bench (Airline) 76.7% 70.7% 40% 72.7%
Tau3-Bench (Banking) 38.1% 10.6% 5.7% 15.5%
BFCL v4 (overall) 61.4% 67.2% 58% 61%

Coding and software engineering

Reported coding results are 85.9 on LiveCodeBench v6, 66.4 on SWE-bench Verified, and 27.7 on Terminal-Bench 2.1.

Raw code generation looks strong, agentic software engineering looks ordinary, and the Terminal-Bench number says long multi-step terminal work is not what this model is for.

There is a contamination caveat you should read before quoting any of it.

The technical report notes that 48% of the English and 47% of the German HumanEval evaluation data appeared in training-related material, which inflates HumanEval-style scores.

Several of the suites in the comparison tables are proprietary or vendor-created, which makes the full set difficult to audit, and no verified like-for-like comparison against the current Claude, GPT, or Gemini flagships existed at the time of writing.

Throughput and efficiency

The efficiency claim is the one that survives the vendor-reporting caveat best, because it is a property of the architecture rather than a score.

Throughput here means how much text the model can generate per GPU per second. Aleph Alpha measures it in bytes rather than tokens, so models with different tokenizers can be compared fairly.

Kolibri 1 sits at 71% average on Aleph Alpha's German benchmark suite while decoding roughly 47,000 bytes/s per GPU. GPT OSS A5B reaches 70% at about 34,500 bytes/s, Qwen 3.6 A3B 67% at about 38,500, Gemma 4 A4B 66% at about 43,000, and Nemotron Super A12B 68% at about 24,000.

Serving roughly 2x the decoded text per GPU of the Nemotron model at a higher German average is the practical argument for picking Kolibri in a self-hosted stack.

If you are paying for your own GPUs rather than per token, throughput per GPU is the number that shows up on the invoice.

Kolibri 1 Pricing and Availability

There is no published token price for Kolibri 1.

Aleph Alpha has not released a hosted API price list, and no confirmed Kolibri API model ID appeared in the official API documentation as of October 5, 2026.

Enterprise deployment runs through the Aleph Alpha sales team, and the repository identifier Aleph-Alpha/Kolibri-1 should not be treated as an API model ID.

The weights themselves are free under Apache 2.0, so your cost is GPU time.

The FP8 memory footprint is roughly 78 GB, with a stated minimum of 2x A100 80 GB, 2x H100 SXM5, 1x H200, 1x B200, or 1x B300, and a recommended configuration of 2x H100 SXM5, 2x H200, 1x B200, or 1x B300. The model is generally available, not a preview, with no waitlist or region lock.

For a sense of what you are trading against, our GPT-6.1 Sol coverage puts that model at $2 input / $10 output per million tokens. A third-party pricing summary lists the older GPT-5.6 Sol at $5 / $30 and GPT-5.6 Luna at $0.20 / $1.20 after OpenAI's July 30 price cut.

Self-hosting wins on unit economics only once your volume clears the fixed cost of a B200.

How to Get Access to Kolibri 1?

Kolibri 1 is open-weight under Apache 2.0, which permits commercial use, modification, and redistribution with no user or revenue thresholds.

The weights live at Aleph-Alpha/Kolibri-1 on Hugging Face in float8_e4m3fn. The model card documents serving through vLLM only, using Aleph Alpha's aleph-alpha-inference plugin. Community GGUF and MLX builds appeared within days, but Aleph Alpha has not validated them, and I couldn't find an official Ollama entry.

For a served deployment, 1x H200 or 1x B200 holds the ~78 GB of FP8 weights on a single card. Use --tensor-parallel-size 2 on H100s.

Keep --max-model-len at or below 262,144 unless you have specifically tested longer contexts. Going beyond that needs an extra --hf-overrides flag, which the model card documents.

Install the plugin and start the server:

pip install 'aleph-alpha-inference>=1'

# Add --tensor-parallel-size 2 on 2x H100
vllm serve Aleph-Alpha/Kolibri-1 \
  --kv-cache-dtype fp8 \
  --max-model-len 262144 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

Then query it through the OpenAI-compatible API, using the sampling parameters Aleph Alpha recommends (temperature 1.0, top_p 0.97, top_k 128):

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="Aleph-Alpha/Kolibri-1",
    messages=[{
        "role": "user",
        "content": "Fasse die wichtigsten Fristen aus diesem Bescheid in drei Punkten zusammen.",
    }],
    temperature=1.0,
    top_p=0.97,
    max_tokens=512,
    extra_body={
        "top_k": 128,
        "chat_template_kwargs": {"reasoning_effort": "medium"},
    },
)
print(response.choices[0].message.content)

Reasoning effort accepts low, medium, and high, and you can switch thinking off entirely if latency matters more than depth.

Outputs will vary with sampling parameters, and the model card notes that they can diverge slightly even with sampling disabled.

For deeper setup work on self-hosted inference and agent loops, our API tutorial on building a migration agent covers the tool-permission and steering patterns that transfer directly.

Kolibri 1 and the Sovereignty Question: The Cohere Merger

Kolibri one is notionally a 'sovereign AI', meaning a country or organization can run, audit, and control its AI on its own infrastructure and under its own jurisdiction, without depending on a foreign provider's API, pricing, or terms of service. For Kolibri 1, that case rests on Apache 2.0 weights you can run air-gapped (disconnected from the internet), European training infrastructure, and EU AI Act documentation.

However, two pieces of corporate context sit behind this launch.

Aleph Alpha has signed a binding merger agreement with Canadian company Cohere to form what they describe as the first transatlantic sovereign AI solution.

The combined company will operate under the Cohere name, with dual headquarters in Toronto and Berlin, and Heidelberg continuing as a research center. The deal still needs regulatory approval.

A sovereignty pitch depends on whoever owns the model continuing to support it, and the Aleph Alpha brand is due to disappear into Cohere once the merger closes, so both items are worth tracking alongside the benchmarks.

Final Thoughts

Aleph Alpha is betting that sovereignty, Apache 2.0 licensing, and documented EU AI Act compliance matter more to European public-sector buyers than a few points on MMLU-Pro.

That bet looks reasonable, and the company's merger with Cohere suggests it knows it cannot win on scale alone.

I think it's fair to say that Kolibri 1 is the best German-language open-weight model at this active-parameter size that we've seen documented this well, but it is not the model I'd reach for if I needed long multi-turn agent runs or closed-book factual recall.

Among comparable small-active MoE models, Qwen3.6 35B-A3B still wins that fight. If you can afford a dense 27B model, Qwen3.8 27B wins it outright.

If you want to get up to speed on running and evaluating models like this, start with our AI Fundamentals skill track.

FAQs

Is Kolibri 1 free to use?

The weights are free under an Apache 2.0 license, which permits commercial use, modification, and redistribution with no revenue or user thresholds. There is no published hosted API price and no confirmed Kolibri API model ID as of October 5, 2026, so your cost is the GPU time you pay for. Enterprise deployment runs through Aleph Alpha's sales team.

What hardware do you need to run Kolibri 1?

The FP8 weights have a memory footprint of roughly 78 GB. Aleph Alpha lists a minimum of 2x A100 80 GB, 2x H100 SXM5, 1x H200, 1x B200, or 1x B300, and recommends 2x H100 SXM5, 2x H200, 1x B200, or 1x B300. One user reported around 170 tokens per second at 8-bit quantization on a single workstation GPU.

How does Kolibri 1 compare to Qwen3.6 35B-A3B?

Kolibri 1 leads on competition math and German reasoning, scoring 96% on AIME 2026 (EN) against Qwen3.6's 91% and 90% on AIME 2026 (DE) against 84.4%. Qwen3.6 wins on MMLU-Pro CoT (84.3% vs 80%), the AA-Omniscience Index (42.35% vs 33.6%), and BFCL v4 tool calling (67.2% vs 61.4%). The exception is Tau3-Bench (Banking), where Kolibri scores 38.1% against Qwen3.6's 10.6%.

Where can I download Kolibri 1?

The weights are published on Hugging Face as Aleph-Alpha/Kolibri-1, and an Ollama entry is listed as kolibri. The model was not listed in the OpenRouter or models.dev catalogs as of October 5, 2026, and no hosted inference provider was publicly available at launch.

What is Kolibri 1 best used for?

Aleph Alpha positions it for multi-step reasoning, retrieval-augmented generation, agentic tool calling, coding, and German- and English-language assistant work. It is strongest on math, German-language reasoning, and raw code generation, and weakest on closed-book factual recall, long multi-turn tool use, and agentic software engineering. It is text-only with no image, audio, or video support.

Matt Crabtree's photo
Author
Matt Crabtree
LinkedIn

A senior editor in the AI and edtech space. Committed to exploring data and AI trends.  

Topics
Artificial Intelligence
Large Language Models

Top DataCamp Courses

Course

AI-Assisted Coding for Developers

1 hr 30 min
10.5K
Boost your coding with AI—guide your coding assistant to write, test, and document code effectively.
See DetailsRight Arrow
Start Course
See MoreRight Arrow
Related

blog

Kimi K3: Moonshot AI's Newest and Best Open-Source Model

Read about Kimi K3 : Everything we know about Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-source LLM and the largest open-weight model released to date. See benchmarks, pricing, and API features.
Josef Waples's photo

Josef Waples

11 min

blog

Inkling: Thinking Machines' First Open-Weights Model

Inkling is Thinking Machines' first open-weights model, a 975B-parameter MoE with 1M-token context, native multimodal input, and controllable thinking effort.
Matt Crabtree's photo

Matt Crabtree

10 min

blog

GLM-5.2: Features, Setup, Benchmarks, and Model Switching Guide

Z.ai's GLM-5.2 ships with a 1M token context window, two reasoning effort levels, and free access across all GLM Coding Plan tiers.
Matt Crabtree's photo

Matt Crabtree

11 min

blog

Qwen3.8-Max: Alibaba's 2.4T Model for Coding and Autonomous Work

Qwen3.8-Max is Alibaba's 2.4T-parameter flagship, topping OSWorld-Verified and PaperBench. Here's what's new, the benchmarks, and what it means for the industry.
Matt Crabtree's photo

Matt Crabtree

10 min

Llama 3.2 is now multimodal

blog

Llama 3.2 Guide: How It Works, Use Cases & More

Meta releases Llama 3.2, which features small and medium-sized vision LLMs (11B and 90B) alongside lightweight text-only models (1B and 3B). It also introduces the Llama Stack Distribution.
Alex Olteanu's photo

Alex Olteanu

8 min

Tutorial

Kimi K2 Thinking: Open-Source LLM Guide, Benchmarks, and Tools

Hands-on tutorial to run Kimi K2 Thinking, build tool-calling workflows, view transparent reasoning, and benchmark against GPT-5 and Claude 4.5.
Bexruz (Bex) Tuychiev's photo

Bexruz (Bex) Tuychiev

15 min

See MoreSee More