AI and machine learning blogs — page 2
Current feed archive. Entries come from the loaded publisher feeds and change as those feeds update. Search and saved-story filters apply to this page.
Browse the news
Headlines grouped by publisher. Search this page by source or headline; use the archive for older entries.
Spacing
Order
Teaching Everyone to Fish for Tokens
GLM-5.3: How Chinese labs keep stride with the frontier
I wrote an AI textbook — how long until AI can do it better?
5 useful things you'll learn in my new post-training textbook (shipping now!)
Lessons from the hacks
Introducing our Artifacts Hub and Adoption Dashboard
Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier
Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next
Kimi K3: The open-weights escalation
6 months to live for open models
Becoming Human AI Is Expanding — Here’s What’s Changing
Becoming Human AI Is Expanding — Here’s What’s Changing For years, this has been where you’ve found us — through Medium, whenever we published something worth your time. That’s ch…
What is RAFT? RAG + Fine-Tuning
What is RAFT? RAG + Fine-Tuning In simple terms, retrieval-augmented fine-tuning, or RAFT, is an advanced AI technique in which retrieval-augmented generation is joined with fine-…
Modern Operating Systems for AI Agents
Modern Operating Systems for AI Agents An operating system (OS) is the fundamental software that acts as an intermediary between computer hardware and user applications. It manage…
NLP in 2026: Trends, Use Cases & Future of Language AI | Shaip
NLP in 2026: Trends, Use Cases & Future of Language AI | Shaip Every day, your organization produces a mountain of words. Support tickets, contracts, clinical notes, customer revi…
AI Bias in Hiring Tools: Real Disasters and How Teams Fixed Them
Hiring tools failing women and minorities? Real cases like Amazon's flop. Step-by-step fixes – data tweaks, fairness scores.
Harness Engineering for Self-Improvement
The concept of recursive self-improvement (RSI) dates back to I. J. Good (1965), where he defined an “ultraintelligent machine” as a system that can surpass humans in all intellec…
Scaling Laws, Carefully
Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model…
Context Graphs vs Vector RAG vs Raw Context - Agentic memory benchmark
We benchmark context graphs against other methods of agentic memory, and discuss what each method got right and wrong.
What Are Context Graphs? (And Why AI Agents Need Them)
We explain why AI agents need to capture decisions, how context graphs achieve this, and how AI agents can use these graphs.
SAP Sapphire 2026: The Complete Breakdown
All 25 announcements from Orlando, what's actually shipping vs. what's marketing — and what it means for the document and data foundation of the autonomous enterprise. SAP's Sapph…
Claude for Legal Teams: Contract Review, Compliance and Due Diligence
See how the Claude legal plugin helps in-house legal teams with contract review, compliance scanning, due diligence, obligations tracking, and drafting.
Vibe Coding Best Practices: 5 Claude Code Habits for Better Agentic Coding
Learn 5 practical vibe coding best practices for Claude Code and coding agents: CLAUDE.md, planning, review agents, safer prompts, and diff review.
AI Benchmarks Explained: GPQA, SWE-bench, Chatbot Arena and What They Actually Measure
Learn what MMLU, GPQA Diamond, SWE-bench, HealthBench, and Chatbot Arena actually measure, and how labs game benchmark scores.
Why AI-Native IDP Platforms Outperform ABBYY and Kofax in Modern Document Workflows
Evaluating IDP vendors? Compare Nanonets vs ABBYY and Kofax across architecture, operating model, and TCO to see why AI-native wins for IDP.
Why You Hit Claude Limits So Fast: AI Token Limits Explained
Learn what AI tokens are, why Claude hits limits fast, and how to cut waste from context windows, history, files, tools, and reasoning.
Domain-Specific AI Should Focus on Workflows Rather Than Modeling
I spent the past few years working on AI for health. Starting with custom multimodal encoders, post-training, and sophisticated multi-agent architecturs, I now see modeling work b…
AI Evaluation is Becoming an Exciting Standalone Discipline
Having worked on robustness problems during my PhD, I see many of the characteristics appearing in the evaluation of LLMs and AI systems. Adversarial attacks such as jailbreaks ar…
No matching sources found.