plaintextmlAI research & developer tools

News index / Blogs

AI and machine learning blogs

Read practitioner writing on LLMs, data science, model engineering, product lessons, and applied ML tutorials.

Reading queue

05 entries
05

Laser gems

AI Weirdness · Blogs · 6d ago

Browse the news

Latest from each source

Headlines grouped by publisher. Search this page by source or headline; use the archive for older entries.

Spacing
Order
Blogs Hugging Face - Blog

20 entries on this page

Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community 9h ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem 19h ago A Blog post by Multiverse Computing on Hugging Face tokenizers v1: encode, decode and scaling, measured 1d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Your Agent Aced the Task. Will It Do It Again? 6d ago A Blog post by IBM Research on Hugging Face Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL 12d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Rebuilding AUTOMATIC1111 with Gradio Workflow 12d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license 12d ago A Blog post by IBM Research on Hugging Face Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic 13d ago A Blog post by Multiverse Computing on Hugging Face NeoMME: an efficient Multimodal-native and Multilingual Encoder 18d ago A Blog post by H company on Hugging Face Give Your Coding Agents a Memory You Own 19d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps 19d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Training a coding model to paint watercolours with TRL and OpenEnv 19d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Real-Time Intelligence with IBM Time Series Models on Confluent 19d ago A Blog post by IBM Research on Hugging Face BenchMIRT: What are LLM benchmarks actually measuring? 20d ago A Blog post by Ai2 on Hugging Face Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI 21d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. The Open ASR Leaderboard Adds Its First Global South Language 25d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers 27d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Granite 4.2 LLMs: How They're Built 27d ago A Blog post by IBM Granite on Hugging Face Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC 27d ago A Blog post by IBM Granite on Hugging Face Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original 27d ago A Blog post by Multiverse Computing on Hugging Face
Blogs John D. Cook

20 entries on this page

Blogs Interconnects AI

20 entries on this page

Blogs One Useful Thing

20 entries on this page

Blogs AI Weirdness

8 entries on this page

Blogs Towards Data Science

20 entries on this page

Your Model Isn't Done Until Someone Else Can Call It 8d ago Your AI Adoption Lift Is a Selection Effect 8d ago One Capital Letter Was Silently Breaking My AI Support Bot, and It Wasn't in the New Model 9d ago Stop Managing Alarms: An Incident-First Blueprint for Telecom AIOps 9d ago What large operators can teach us about turning alert fatigue into faster, safer service assurance Coding Agents Don't Need Longer History — They Need Intent Continuity 10d ago Software Design in the Age of AI 10d ago The 95% Illusion: Why Your Confidence Interval Isn't What You Think It Is 10d ago Demystifying Anthropic's J-Space: A Mathematical Primer 10d ago How to 5x Your Communication Effectiveness with Claude Code 11d ago Better understand the intent of your coding agents What SHAP Can't Explain About Agentic AI Fraud 11d ago Optimizing LLM Inference Costs in Multi-Agent Systems with Adaptive Model Routing 11d ago Moving from static model assignment to intelligent, task-level LLM selection. Who Questions What Works: When Should We Retest Our Assumptions? 11d ago A model is only as reliable as the assumptions behind it Getting started with dbt 12d ago The Symmetry That Breaks Neural Network Averaging 12d ago When One Process Becomes Too Much: Splitting a Pipeline into MCP Services 12d ago 10 Statistical Traps We Often Overlook 12d ago How to Maximize GPT-6 Astra 13d ago The Model Validation Playbook for GenAI: Lessons from Banking 13d ago A Beginner’s Guide to World Models 13d ago Introducing ShipAI 13d ago Towards Data Science launches a video showcase for real-world AI work
Blogs Ahead of AI

20 entries on this page

Blogs Salmon Run

20 entries on this page

Book Review: Data Centric Machine Learning with Python 23d ago I recently read the book Data Centric Machine Learning with Python . The thesis of the book is that Machine Learning (ML) pipelines benefit ... Book Review: Domain Specific Small Language Models 64d ago Artificial Intelligence (AI) powered applications are changing the way we consume and use information. Retrieval Augmented Generation (RAG) ... Book Review: Software Engineering for Data Scientists 232d ago As a Software Engineer (backend Web Development then Search) turned Data Scientist, I was particularly interested in what the book Software ... Book Review: Transformers In Action 254d ago The Attention Is All You Need paper proposed the Transformer Architecrture as an improvement to the dominant encoder-decoder models of the ... Trip Report: PyData Global 2025 269d ago I attended PyData Global 2025 earlier this month. I had hoped to write this up earlier, but I've been busy, so only now getting the time Ch... Book Review: Time Series Forecasting using Foundation Models 344d ago As someone who primarily works in NLP and Search in the Health Domain, I don't have much use for Time Series. However, while exploring the F... Book Review: Statistics every Programmer Needs 366d ago I recently read Statistics every Programmer Needs by Gary Sutton. I am probably a good target audience for the book since I used to be a so... Book Review: Hands-On Artificial Intelligence for IoT 450d ago For those in similar professional circles as I am in, i.e. looking forward into the Generative AI space, yet with one foot pragmatically and... Book Review: Essential Graph RAG 463d ago Coming from a background of Knowledge Graph (KG) backed Medical Search, I don't need to be convinced about the importance of manually curate... Packaging ML Pipelines from Experiment to Deployment 629d ago As an ML Engineer, we are generally tasked with solving some business problem with technology. Typically it involves leveraging data assets ... Trip Report - PyData Global 2024 652d ago I attended PyData Global 2024 last week. Its a virtual conference, so I was able to attend it from the comfort of my home, although presenta... Using Knowledge Graphs to enhance Retrieval Augmented Generation 716d ago Retrieval Augmented Generation (RAG) has become a popular approach to harness LLMs for question answering using your own corpus of data. Typ... Experiments with Prompt Compression 784d ago I recently came across Prompt Compression (in the context of Prompt Engineering on Large Language Models) on this short course on Prompt Com... Table Extraction from PDFs using Multimodal (Vision) LLMs 813d ago Couple of weeks ago a colleague and I participated in an internal hackathon where the task was to come up with an interesting use case using... Book Report: Pandas Workout 820d ago Unlike many Data Scientists, I didn't automatically reach for Pandas when I needed to analyze data. I came upon this discipline (Data Scien... Finetuning RAGAS Metrics using DSPy 856d ago Last month, I decided to sign-up for the Google AI Hackathon , where Google provided access to their Gemini Large Language Model (LLM) and ... Performance Analysis of Float vs Byte vs Binary Vectors on OpenSearch 860d ago I've been working on an application where, given an input string, the objective is to recommend an output string that is similar to the inpu... KGC/HCLS 2024 Trip Report 867d ago I was at KGC (Knowledge Graph Conference) 2024 , which is happening May 6-10 at Cornell Tech . I was presenting (virtually) at their Health ... Book Report: Machine Learning for Drug Discovery 912d ago Drug Discovery is a field where biochemists (and more recently computer scientists) turn ideas into potential medications. I first came acro... Hierarchical (and other) Indexes using LlamaIndex for RAG Content Enrichment 918d ago At our weekly This Week in Machine Learning (TWIML) meetings, (our leader and facilitataor) Darin Plutchok pointed out a LinkedIn blog post ...
Blogs Becoming Human: Artificial Intelligence Magazine - Medium

10 entries on this page

Becoming Human AI Is Expanding — Here’s What’s Changing 39d ago Becoming Human AI Is Expanding — Here’s What’s Changing For years, this has been where you’ve found us — through Medium, whenever we published something worth your time. That’s ch… What is RAFT? RAG + Fine-Tuning 40d ago What is RAFT? RAG + Fine-Tuning In simple terms, retrieval-augmented fine-tuning, or RAFT, is an advanced AI technique in which retrieval-augmented generation is joined with fine-… Modern Operating Systems for AI Agents 40d ago Modern Operating Systems for AI Agents An operating system (OS) is the fundamental software that acts as an intermediary between computer hardware and user applications. It manage… NLP in 2026: Trends, Use Cases & Future of Language AI | Shaip 40d ago NLP in 2026: Trends, Use Cases & Future of Language AI | Shaip Every day, your organization produces a mountain of words. Support tickets, contracts, clinical notes, customer revi… AI Bias in Hiring Tools: Real Disasters and How Teams Fixed Them 132d ago Hiring tools failing women and minorities? Real cases like Amazon's flop. Step-by-step fixes – data tweaks, fairness scores. AGI in 2025 |Do you think what matters today will still matter in the coming months? TL;DR: No! 595d ago When Algorithms Dream of Photons: Can AI Redefine Reality Like Einstein? 595d ago Be Part of the AI Revolution at the Chatbot Conference Tomorrow! 728d ago Join the Most-Awaited Chatbot Conference 731d ago Limited Time Offer: Get Your Exclusive Online Passes to the Chatbot Conference — Act Fast! 732d ago
Blogs Lil'Log

20 entries on this page

Harness Engineering for Self-Improvement 80d ago The concept of recursive self-improvement (RSI) dates back to I. J. Good (1965), where he defined an “ultraintelligent machine” as a system that can surpass humans in all intellec… Scaling Laws, Carefully 90d ago Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model… Why We Think 509d ago Special thanks to John Schulman for a lot of super valuable feedback and direct edits on this post. Test time compute (Graves et al. 2016, Ling, et al. 2017, Cobbe et al. 2021) an… Reward Hacking in Reinforcement Learning 663d ago Reward hacking occurs when a reinforcement learning (RL) agent exploits flaws or ambiguities in the reward function to achieve high rewards, without genuinely learning or completi… Extrinsic Hallucinations in LLMs 807d ago Hallucination in large language models usually refers to the model generating unfaithful, fabricated, inconsistent, or nonsensical content. As a term, hallucination has been somew… Diffusion Models for Video Generation 893d ago Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation.… Thinking about High-Quality Human Data 960d ago [Special thank you to Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper “Vox populi”) and nice feedback. 🙏 ] High-quality data is the fuel for modern dat… Adversarial Attacks on LLMs 1063d ago The use of large language models in the real world has strongly accelerated by the launch of ChatGPT. We (including my team at OpenAI, shoutout to them) have invested a lot of eff… LLM Powered Autonomous Agents 1187d ago Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT, GPT-Engineer and BabyAGI, serve as insp… Prompt Engineering 1287d ago Prompt Engineering, also known as In-Context Prompting, refers to methods for how to communicate with LLM to steer its behavior for desired outcomes without updating the model wei… The Transformer Family Version 2.0 1334d ago Many new Transformer architecture improvements have been proposed since my last post on “The Transformer Family” about three years ago. Here I did a big refactoring and enrichment… Large Transformer Model Inference Optimization 1350d ago [Updated on 2023-01-24: add a small section on Distillation.] Large transformer models are mainstream nowadays, creating SoTA results for a variety of tasks. They are powerful but… Some Math behind Neural Tangent Kernel 1474d ago Neural networks are well known to be over-parameterized and can often easily fit data with near-zero training loss with decent generalization performance on test dataset. Although… Generalized Visual Language Models 1565d ago Processing images to generate text, such as image captioning and visual question-answering, has been studied for years. Traditionally such systems rely on an object detection netw… Learning with not Enough Data Part 3: Data Generation 1620d ago Here comes the Part 3 on learning with not enough data (Previous: Part 1 and Part 2). Let’s consider two approaches for generating synthetic data for training. Augmented data. Giv… Learning with not Enough Data Part 2: Active Learning 1675d ago This is part 2 of what to do when facing a limited amount of labeled data for supervised learning tasks. This time we will get some amount of human labeling work involved, but wit… Learning with not Enough Data Part 1: Semi-Supervised Learning 1752d ago When facing a limited amount of labeled data for supervised learning tasks, four approaches are commonly discussed. How to Train Really Large Models on Many GPUs? 1824d ago [Updated on 2022-03-13: add expert choice routing.] [Updated on 2022-06-10]: Greg and I wrote a shorted and upgraded version of this post, published on OpenAI Blog: “Techniques fo… What are Diffusion Models? 1899d ago [Updated on 2021-09-19: Highly recommend this blog post on score-based generative modeling by Yang Song (author of several key papers in the references)]. [Updated on 2022-08-27:… Contrastive Representation Learning 1940d ago The goal of contrastive representation learning is to learn such an embedding space in which similar sample pairs stay close to each other while dissimilar ones are far apart. Con…
Blogs Nanonets Blog | AI Agents for Enterprise Data Processing

15 entries on this page

Context Graphs vs Vector RAG vs Raw Context - Agentic memory benchmark 82d ago We benchmark context graphs against other methods of agent memory, and discuss their benefits. What Are Context Graphs? (And Why AI Agents Need Them) 82d ago A context graph is the part of agent memory that captures decisions. We explain why AI agents need to capture decisions, how context graphs capture them, and how agents use them. SAP Sapphire 2026: The Complete Breakdown 123d ago All 25 announcements from Orlando, what's actually shipping vs. what's marketing — and what it means for the document and data foundation of the autonomous enterprise. SAP's Sapph… Claude for Legal Teams: Contract Review, Compliance and Due Diligence 153d ago See how the Claude legal plugin helps in-house legal teams with contract review, compliance scanning, due diligence, obligations tracking, and drafting. Vibe Coding Best Practices: 5 Claude Code Habits for Better Agentic Coding 158d ago Learn 5 practical vibe coding best practices for Claude Code and coding agents: CLAUDE.md, planning, review agents, safer prompts, and diff review. AI Benchmarks Explained: GPQA, SWE-bench, Chatbot Arena and What They Actually Measure 164d ago Learn what MMLU, GPQA Diamond, SWE-bench, HealthBench, and Chatbot Arena actually measure, and how labs game benchmark scores. Why AI-Native IDP Platforms Outperform ABBYY and Kofax in Modern Document Workflows 164d ago Evaluating IDP vendors? Compare Nanonets vs ABBYY and Kofax across architecture, operating model, and TCO to see why AI-native wins for IDP. Why You Hit Claude Limits So Fast: AI Token Limits Explained 166d ago Learn what AI tokens are, why Claude hits limits fast, and how to cut waste from context windows, history, files, tools, and reasoning. Did Google's TurboQuant Actually Solve AI Memory Crunch? 172d ago Google’s TurboQuant promises 6x KV-cache compression. Here’s what it means for AI memory, HBM demand, and the broader memory crunch. Claude for Finance Teams: Investment Banking, DCF Models, Reconciliation & Variance Analysis 182d ago See how finance teams use Claude for one-pagers, CIMs, comps, DCF models, reconciliations, and variance commentary, plus key human checks. AI Agent Hacks McKinsey: 5 Situations When You Should Not Deploy Agents 191d ago McKinsey hacked in 2 hours. 5 situations where AI agents will fail. Production permissions, regulated data, legacy systems—check before deploy. Are OpenAI and Google intentionally downgrading their models? 194d ago Yes, OpenAI and Google degrade their models. OpenAI admitted silent updates after denying it. Gemini redirects models. With full evidence. We ran 16 AI Models on 9,000+ Real Documents. Here's What We Found. 194d ago We benchmarked GPT-5.4, Gemini 3.1 Pro, Claude Opus, Sonnet, and 12 others on 3 Open OCR Benchmarks AI Arms Race Has Real Numbers: Pentagon vs China 2026 199d ago AI targeting: 900 strikes in 12 hours. Anthropic banned from Pentagon mid-operation, OpenAI switched in. China building parallel systems. Stop Paying for AI You Don't Use: The Case for Fine-Tuned Models 203d ago Processing 10K documents daily costs $50K/year with GPT. Fine-tuned models: $5K with better latency and stable accuracy. Make the right choice.
Blogs David Stutz

2 entries on this page

Blogs Damian Bogunowicz - dtransposed

10 entries on this page

No matching sources found.