Category
Technologies
LLM Articles
Keep up to date with the latest techniques, tools, and research in Large Language Models. Our blog talks about data science, uses, & responsible AI practices.
Other technologies:
Training 2 or more people?Try DataCamp for Business
GPT-6 Astra vs Claude Fable 5.1: Performance, Pricing, and Which to Use
Two frontier models arrived at exactly the same list price, and OpenAI's own benchmark table and the independent index disagree on which one leads.
Tom Farnschläder
September 5, 2026
GPT-6 Astra: Features, Benchmarks, Pricing, and How to Access It
OpenAI's GPT-6 Astra tops computer use, coding, and math benchmarks. Full breakdown of features, scores vs Claude and Gemini, pricing, and how to access it.
Matt Crabtree
September 3, 2026
Muse Spark 1.3: Meta's Agentic and Coding Model Update
Meta's Muse Spark 1.3 improves agentic workflows and coding, using ~20% fewer tool calls and ~25% fewer tokens than 1.2, with a 1M context window.
Matt Crabtree
September 3, 2026
Gemini 3.8 Flash and 3.8 Flash Cyber: Features, Benchmarks, and Pricing
Google's third Flash release in six weeks pushes coding and agentic reasoning at the same low price as 3.7 Flash, plus a dedicated cybersecurity variant.
Matt Crabtree
September 2, 2026
Claude Fable 5.1: Features, Benchmarks, and Pricing
Fable 5.1 holds Fable 5's list price but cuts cache reads 75%, and it now leads Claude Opus 5 on every benchmark Anthropic published.
Tom Farnschläder
September 2, 2026
GLM-5.3-Flash vs Qwen3.8-Flash-Next: Coding, Cost, and Access
Z.ai's and Alibaba's newest models compete for the same segment. GLM-5.3-Flash leads shared coding benches and is live; Qwen3.8-Flash-Next is the lightweight Qwen4 preview with a queued API.
Tom Farnschläder
August 31, 2026
Grok 4.6 vs Claude Opus 5: Coding, Price, and Which to Use
Grok 4.6 undercuts Claude Opus 5 at $2/$6 versus $5/$25, but Opus 5 still leads knowledge-work Elo and our aurora shader test. Here's who should use which.
Tom Farnschläder
August 30, 2026
GLM-5.3-Flash: Features, Benchmarks, Pricing, and How It Compares
Z.ai's cost-optimized GLM-5.3-Flash (Ox Alpha) lands near-frontier coding and agentic scores at roughly a tenth of the price of GLM-5.3.
Matt Crabtree
August 27, 2026
Grok 4.6 vs GPT-5.6 Sol: Benchmarks, Pricing, and a Hands-On Test
Grok 4.6 and GPT-5.6 Sol score the same 61 on the Artificial Analysis Intelligence Index, so the real decision is about cost shape, context size, and turns per task.
Tom Farnschläder
August 27, 2026
Qwen3.8-Flash-Next: Alibaba's Cost-Efficient Preview of Qwen4
Qwen3.8-Flash-Next is Alibaba's open-weight 125B MoE model previewing the Qwen4 architecture. It beats Claude Opus 4.6 Max on most coding and agent benchmarks.
Matt Crabtree
August 27, 2026
Grok 4.6: Features, Benchmarks, Pricing, and Comparisons
SpaceXAI's new model, Grok 4.6, matches GPT-5.6 Sol's Intelligence Index score at a lower measured price. See benchmarks, agent features, pricing, and how it compares with Grok 4.5 and Claude Sonnet 5.
Khalid Abdelaty
August 21, 2026
Gemini 3.7 Flash: Features, Benchmarks, and Pricing
Google's Gemini 3.7 Flash targets coding and agentic workflows at half the launch price of 3.6 Flash. Here's what's new, the benchmarks, and where it fits.
Matt Crabtree
August 14, 2026