Loading…

AI & Machine Learning News Hub

Research, releases, and applied work in AI & ML

What's New

Top 5 Across All Sources
  1. Corrupted apostrophes

    John D. Cook · 4h ago
  2. How not to calculate cosine

    John D. Cook · 14h ago
Latest
John D. CookCorrupted apostrophesHugging Face - BlogTutorMoments: Do AI tutors know when to help and when to hold back?Towards Data ScienceMatplotlib vs Plotly: Which Python Chart Tool Should You Choose?John D. CookHow not to calculate cosineTowards Data ScienceLoop Engineering for Listing Questions: When the Answer Is Every Passage, Not the Top OneTowards Data ScienceThe Problem with pandas Isn’t Performance. It’s Cognitive Overhead.John D. Cookcos(200!)Towards Data ScienceMy Fall-Detection Model Scored 94%, and It Was Lying to MeTowards Data ScienceI Built an AI Data Agent Which Can Query Data and Answer Business Questions. Here’s How.Towards Data ScienceLast Month’s Machine Learning Lessons LearnedTowards Data ScienceI Built a Tool-Calling Agent in Python. Here’s How I Debugged ItJohn D. CookCalculating log(1000!)Towards Data ScienceLoop Engineering for Cross-References: When RAG Answers ‘see Section 7.2’ Instead of the Actual AnswerHugging Face - BlogBaseten on Hugging Face Inference Providers 🔥John D. CookThe code that didn’t breakTowards Data ScienceHow a Frontier Model Gets Built, Read from the Kimi K3 ReportTowards Data ScienceIntroduction to Semi-Supervised LearningJohn D. CookEnumerating trees and circlesTowards Data ScienceIs This Slop? Detecting AI-Generated Content Without a ModelTowards Data ScienceBuilding Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAGJohn D. CookCorrupted apostrophesHugging Face - BlogTutorMoments: Do AI tutors know when to help and when to hold back?Towards Data ScienceMatplotlib vs Plotly: Which Python Chart Tool Should You Choose?John D. CookHow not to calculate cosineTowards Data ScienceLoop Engineering for Listing Questions: When the Answer Is Every Passage, Not the Top OneTowards Data ScienceThe Problem with pandas Isn’t Performance. It’s Cognitive Overhead.John D. Cookcos(200!)Towards Data ScienceMy Fall-Detection Model Scored 94%, and It Was Lying to MeTowards Data ScienceI Built an AI Data Agent Which Can Query Data and Answer Business Questions. Here’s How.Towards Data ScienceLast Month’s Machine Learning Lessons LearnedTowards Data ScienceI Built a Tool-Calling Agent in Python. Here’s How I Debugged ItJohn D. CookCalculating log(1000!)Towards Data ScienceLoop Engineering for Cross-References: When RAG Answers ‘see Section 7.2’ Instead of the Actual AnswerHugging Face - BlogBaseten on Hugging Face Inference Providers 🔥John D. CookThe code that didn’t breakTowards Data ScienceHow a Frontier Model Gets Built, Read from the Kimi K3 ReportTowards Data ScienceIntroduction to Semi-Supervised LearningJohn D. CookEnumerating trees and circlesTowards Data ScienceIs This Slop? Detecting AI-Generated Content Without a ModelTowards Data ScienceBuilding Document Structure with Loop Engineering: Recovering a PDF’s Outline from Body Typography for RAG

By Source

Feeds organized so you can skim by site.

Density Sort
JD
John D. Cook
4h ago · 20 items
20 loaded
HF
Hugging Face - Blog
12h ago · 20 items
TutorMoments: Do AI tutors know when to help and when to hold back? 12h ago A Blog post by Ai2 on Hugging Face Baseten on Hugging Face Inference Providers 🔥 2d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Deploy local agents everywhere with LFM2.5-2.6B 3d ago A Blog post by Liquid AI on Hugging Face GPU Management: Why Idle GPUs Are the New Grounded Aircraft 8d ago A Blog post by Dharma-AI on Hugging Face NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics 11d ago A Blog post by NVIDIA on Hugging Face Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident 12d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Bringing Nunchaku 4-bit Diffusion Inference to Diffusers 16d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Grabette: an open system to record robot-manipulation data 18d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science. Newer Models, Same Advantage 22d ago A Blog post by Dharma-AI on Hugging Face Security incident disclosure — July 2026 23d ago We’re on a journey to advance and democratize artificial intelligence through open source and open science.
20 loaded
TD
Towards Data Science
13h ago · 20 items
20 loaded
IA
Interconnects AI
4d ago · 20 items
20 loaded
OU
One Useful Thing
15d ago · 20 items
20 loaded
SR
Salmon Run
19d ago · 20 items
Book Review: Domain Specific Small Language Models 19d ago Artificial Intelligence (AI) powered applications are changing the way we consume and use information. Retrieval Augmented Generation (RAG) ... Book Review: Software Engineering for Data Scientists 187d ago As a Software Engineer (backend Web Development then Search) turned Data Scientist, I was particularly interested in what the book Software ... Book Review: Transformers In Action 209d ago The Attention Is All You Need paper proposed the Transformer Architecrture as an improvement to the dominant encoder-decoder models of the ... Trip Report: PyData Global 2025 224d ago I attended PyData Global 2025 earlier this month. I had hoped to write this up earlier, but I've been busy, so only now getting the time Ch... Book Review: Time Series Forecasting using Foundation Models 299d ago As someone who primarily works in NLP and Search in the Health Domain, I don't have much use for Time Series. However, while exploring the F... Book Review: Statistics every Programmer Needs 321d ago I recently read Statistics every Programmer Needs by Gary Sutton. I am probably a good target audience for the book since I used to be a so... Book Review: Hands-On Artificial Intelligence for IoT 405d ago For those in similar professional circles as I am in, i.e. looking forward into the Generative AI space, yet with one foot pragmatically and... Book Review: Essential Graph RAG 418d ago Coming from a background of Knowledge Graph (KG) backed Medical Search, I don't need to be convinced about the importance of manually curate... Packaging ML Pipelines from Experiment to Deployment 584d ago As an ML Engineer, we are generally tasked with solving some business problem with technology. Typically it involves leveraging data assets ... Trip Report - PyData Global 2024 606d ago I attended PyData Global 2024 last week. Its a virtual conference, so I was able to attend it from the comfort of my home, although presenta...
20 loaded
AO
Ahead of AI
20d ago · 20 items
20 loaded
LL
Lil'Log
35d ago · 20 items
Harness Engineering for Self-Improvement 35d ago The concept of recursive self-improvement (RSI) dates back to I. J. Good (1965), where he defined an “ultraintelligent machine” as a system that can surpass humans in all intellectual activities and design better machines to improve itself.... Scaling Laws, Carefully 45d ago Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss $L$ decreases predictably as we scale up model size $N$, dataset size $D$, and compute $C$, following a powe... Why We Think 464d ago Special thanks to John Schulman for a lot of super valuable feedback and direct edits on this post. Test time compute (Graves et al. 2016, Ling, et al. 2017, Cobbe et al. 2021) and Chain-of-thought (CoT) (Wei et al. 2022, Nye et al. 2021), ... Reward Hacking in Reinforcement Learning 618d ago Reward hacking occurs when a reinforcement learning (RL) agent exploits flaws or ambiguities in the reward function to achieve high rewards, without genuinely learning or completing the intended task. Reward hacking exists because RL enviro... Extrinsic Hallucinations in LLMs 762d ago Hallucination in large language models usually refers to the model generating unfaithful, fabricated, inconsistent, or nonsensical content. As a term, hallucination has been somewhat generalized to cases when the model makes mistakes. Here,... Diffusion Models for Video Generation 848d ago Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation. The task itself is a superset of the image case, since an ima... Thinking about High-Quality Human Data 915d ago [Special thank you to Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper “Vox populi”) and nice feedback. 🙏 ] High-quality data is the fuel for modern data deep learning model training. Most of the task-specific lab... Adversarial Attacks on LLMs 1018d ago The use of large language models in the real world has strongly accelerated by the launch of ChatGPT. We (including my team at OpenAI, shoutout to them) have invested a lot of effort to build default safe behavior into the model during the ... LLM Powered Autonomous Agents 1142d ago Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT, GPT-Engineer and BabyAGI, serve as inspiring examples. The potentiality of LLM extends beyond genera... Prompt Engineering 1242d ago Prompt Engineering, also known as In-Context Prompting, refers to methods for how to communicate with LLM to steer its behavior for desired outcomes without updating the model weights. It is an empirical science and the effect of prompt eng...
20 loaded
Context Graphs vs Vector RAG vs Raw Context - Agentic memory benchmark 37d ago We benchmark context graphs against other methods of agent memory, and discuss their benefits. What Are Context Graphs? (And Why AI Agents Need Them) 37d ago A context graph is the part of agent memory that captures decisions. We explain why AI agents need to capture decisions, how context graphs capture them, and how agents use them. SAP Sapphire 2026: The Complete Breakdown 78d ago All 25 announcements from Orlando, what's actually shipping vs. what's marketing — and what it means for the document and data foundation of the autonomous enterprise. SAP's Sapphire 2026 in Orlando was the most AI-dense keynote in the comp... Claude for Legal Teams: Contract Review, Compliance and Due Diligence 108d ago See how the Claude legal plugin helps in-house legal teams with contract review, compliance scanning, due diligence, obligations tracking, and drafting. Vibe Coding Best Practices: 5 Claude Code Habits for Better Agentic Coding 113d ago Learn 5 practical vibe coding best practices for Claude Code and coding agents: CLAUDE.md, planning, review agents, safer prompts, and diff review. AI Benchmarks Explained: GPQA, SWE-bench, Chatbot Arena and What They Actually Measure 119d ago Learn what MMLU, GPQA Diamond, SWE-bench, HealthBench, and Chatbot Arena actually measure, and how labs game benchmark scores. Why AI-Native IDP Platforms Outperform ABBYY and Kofax in Modern Document Workflows 119d ago Evaluating IDP vendors? Compare Nanonets vs ABBYY and Kofax across architecture, operating model, and TCO to see why AI-native wins for IDP. Why You Hit Claude Limits So Fast: AI Token Limits Explained 121d ago Learn what AI tokens are, why Claude hits limits fast, and how to cut waste from context windows, history, files, tools, and reasoning. Did Google's TurboQuant Actually Solve AI Memory Crunch? 127d ago Google’s TurboQuant promises 6x KV-cache compression. Here’s what it means for AI memory, HBM demand, and the broader memory crunch. Claude for Finance Teams: Investment Banking, DCF Models, Reconciliation & Variance Analysis 137d ago See how finance teams use Claude for one-pagers, CIMs, comps, DCF models, reconciliations, and variance commentary, plus key human checks.
15 loaded
AW
AI Weirdness
45d ago · 15 items
It's 11:00 pm. Do you know where your AI agent is? 45d ago AI agents that email people and post on other people's websites are cursed and we shouldn't make them. Get working on your April Fools Eiffel Tower 128d ago Elevator Surprise: Place a tiny camera in the elevator, and when someone gets in, snap a photo saying, "Welcome to Space Station!" Or build a miniature model of the Eiffel Tower next to it for a dramatic effect. Tower of Pancakes: Create a ... Bonus: More April Fools pranks from Eiffel Tower Llama 128d ago AI Weirdness: the strange side of machine learning When a chatbot runs your store 231d ago You may have heard of people hooking up chatbots to controls that do real things. The controls might run internet searches, run commands to open and read documents and spreadsheets, or even edit or delete entire databases. Whether this soun... Bonus: Incorrect Christmas Carols 232d ago AI Weirdness: the strange side of machine learning Tiny neural net Halloween costumes are the best 283d ago I've been experimenting with getting a tiny circa-2015 recurrent neural network to generate Halloween costumes. Running on a single cat hair-covered laptop, char-rnn has no internet training, but learns from scratch to imitate the data I gi... More tiny neural net costumes 283d ago AI Weirdness: the strange side of machine learning Halloween costumes by tiny neural net 294d ago I've recently been experimenting with one of my favorite old-school neural networks, a tiny program that runs on my laptop and knows only about the data I give it. Without internet training, char-rnn doesn't have outside references to draw ... Bonus: more halloween costumes from tiny neural net 294d ago AI Weirdness: the strange side of machine learning Botober 2025: Terrible recipes from a tiny neural net 311d ago After seeing generated text evolve from the days of tiny neural networks to today's ChatGPT-style large language models, I have to conclude: there's something special about the tiny guys. Maybe it's the way the tiny neural networks string t...
15 loaded
EY
Eugene Yan
48d ago · 20 items
20 loaded
DS
David Stutz
54d ago · 10 items
Domain-Specific AI Should Focus on Workflows Rather Than Modeling 54d ago I spent the past few years working on AI for health. Starting with custom multimodal encoders, post-training, and sophisticated multi-agent architecturs, I now see modeling work becoming less and less important for domain-specific applicati... AI Evaluation is Becoming an Exciting Standalone Discipline 83d ago Having worked on robustness problems during my PhD, I see many of the characteristics appearing in the evaluation of LLMs and AI systems. Adversarial attacks such as jailbreaks are becoming more relevant, edge cases finally become relevant,... RAISE 2025 Panel Statement on Aligning AI to Clinical Values 307d ago Recently, I attended the Responsible AI for Social and Ethical Healthcare 2025 “2.0” Symposium organized by, among others, Harvard Medical School. The symposium featured various panels on topics surrounding generative AI, in particular mult... Some Lessons on Reviews and Rebuttals 551d ago Writing and responding to reviews is the bread and butter of any academic and especially in AI research, PhD students are confronted with both rather early compared to other displicines. Unfortunately, I found that drafting reviews and rebu... Thoughts on Watermarking AI-Generated Content 569d ago Watermarking AI-generated content has the potential to address various problems that generative AI threatens to aggravate — misinformation, impersonation, copyright infringement, web pollution, etc. However, it is also controversial with ma... Thoughts and Lessons for Planning Rater Studies in AI 578d ago With the goal of deploying generative AI systems, rater studies are becoming increasingly common and important. This means more and more researchers and engineers face the challenge of actually planning and conducting rater studies for AI s... Open-Sourcing Relabeled MedQA and Dermatology DDx Datasets 633d ago Dealing with rater disagreement is becoming more important in AI, especially for LLMs and in specialized domains such as health. In the past year, I helped open source two datasets allowing to study rater disagreement in the health domain: ... Thinking About Research Ideas vs. Technology 635d ago In this article, I want to share some thoughts on the difference between research ideas and technology, particularly in machine learning. This distinction is have been contemplating since starting my PhD. After joining Google DeepMind and b... The Importance of Effectively Experimenting in an AI PhD 760d ago Engineering and running experiments are a key component of most PhDs in AI. While there are plenty of more theoretical topics that are often limited to smaller scale experimentation, the trend has definitely been to scale up models, dataset... FAQ for our Monte Carlo Conformal Prediction 817d ago Over the past months, I have given several talks about Monte Carlo conformal prediction and the problem of calibrating with uncertain ground truth, for example, stemming from annotator disagreement. Each time, the audience had great questio...
DB
Damian Bogunowicz - dtransposed
169d ago · 10 items
AK
Andrej Karpathy blog
176d ago · 10 items
DA
Datumbox
468d ago · 20 items
20 loaded
JA
Jay Alammar
500d ago · 10 items
Moving To Substack 500d ago I’m freezing this blog and starting to post on my Substack instead. The authoring experience is much more convenient for me there. Please follow me there, and check out The Illustrated DeepSeek R-1 if you haven’t yet. And check out our How ... Generative AI and AI Product Moats 1187d ago Here are eight observations I’ve shared recently on the Cohere blog and videos that go over them.: Article: What’s the big deal with Generative AI? Is it the future or the present? Article: AI is Eating The World Remaking Old Computer Graphics With AI Image Generation 1315d ago Can AI Image generation tools make re-imagined, higher-resolution versions of old video game graphics? Over the last few days, I used AI image generation to reproduce one of my childhood nightmares. I wrestled with Stable Diffusion, Dall-E ... The Illustrated Stable Diffusion 1404d ago Translations: Chinese, Vietnamese. (V2 Nov 2022: Updated images for more precise description of forward diffusion. A few more images in this version) AI image generation is the most recent AI capability blowing people’s minds (mine included... Applying massive language models in the real world with Cohere 1615d ago A little less than a year ago, I joined the awesome Cohere team. The company trains massive language models (both GPT-like and BERT-like) and offers them as an API (which also supports finetuning). Its founders include Google Brain alums in... The Illustrated Retrieval Transformer 1678d ago Discussion: Discussion Thread for comments, corrections, or any feedback. Translations: Korean, Russian Summary: The latest batch of language models can be much smaller yet achieve GPT-3 like performance by being able to query a database or... Explainable AI Cheat Sheet 1922d ago Introducing the Explainable AI Cheat Sheet, your high-level guide to the set of tools and methods that helps humans understand AI/ML models and their predictions. I introduce the cheat sheet in this brief video: Finding the Words to Say: Hidden State Visualizations for Language Models 2027d ago By visualizing the hidden state between a model's layers, we can get some clues as to the model's Interfaces for Explaining Transformer Language Models 2060d ago Interfaces for exploring transformer language models by looking at input saliency and neuron activation. Explorable #1: Input saliency of a list of countries generated by a language model Tap or hover over the output tokens: Explorable #2: ... How GPT3 Works - Visualizations and Animations 2203d ago Discussions: Hacker News (397 points, 97 comments), Reddit r/MachineLearning (247 points, 27 comments) Translations: German, Korean, Chinese (Simplified), Russian, Turkish The tech world is abuzz with GPT3 hype. Massive language models (lik...
CH
Chip Huyen
569d ago · 10 items
SS
Seita's Place
1149d ago · 10 items
My Faculty Application Experience 1149d ago I spent roughly a year preparing, and then interviewing, for tenure-trackfaculty positions. My job search is finally done, and I am joining theUniversity of ... Books Read in 2022 1315d ago At the end of every year I have a tradition where I write summaries of thebooks that I read throughout the year. Unfortunately this year wasexceptionally bus... Conference on Robot Learning 2022 1319d ago The airplanes on display at the CoRL 2022 banquet. The 2022 Robotics: Science and Systems Conference 1335d ago A photo I took while at RSS 2022 in New York City, on the dinner cruise arranged by the conference. The (In-Person) ICRA 2022 Conference in Philadelphia 1465d ago A photo I took while at ICRA 2022 in Philadelphia. This is the Two New Papers: Learning to Fling and Singulate Fabrics 1471d ago The system for our IROS 2022 paper on singulating layers of cloth with tactile sensing. A Plea to End Harassment 1483d ago Scott Aaronson is a professor of computer science at UT Austin, where hisresearch area is in theoretical computer science. However, he may be more wellknown ... My Paper Reviewing Load 1567d ago UPDATE (July 2026): this post is no longer updated. Please check out my newer paper reviewing document available here: I Stand with Ukraine 1624d ago I stand with Ukraine and firmly oppose Vladimir Putin’s invasion. Books Read in 2021 1680d ago At the end of every year I have a tradition where I write summaries of thebooks that I read throughout the year. Here’s the following post with the roughset ...
MI
ML in Production
1552d ago · 10 items
Driving Experimentation Forward through a Working Group (Experimentation Program Series: Guide 03) 1552d ago We describe how diverse stakeholders can drive experimentation forward through the formation of a working group and what role data science plays. What is an Experimentation program and Who is Involved? (Experimentation Program Series: Guide 02) 1588d ago We define what an experimentation program is and discuss which stakeholder groups should participate in order to drive experimentation forward. Building An Effective Experimentation Program (Experimentation Program Series: Guide 01) 1601d ago An introduction to building an effective experimentation program at your company. Lessons Learned from Writing Online 1644d ago Where I share my story writing MLinProduction and key metrics from building my audience and monetization. Newsletter #087 2075d ago Weekly newsletter dedicated to sharing resources for building and operating production machine learning systems. Newsletter #086 2083d ago Weekly newsletter dedicated to sharing resources for building and operating production machine learning systems. Newsletter #085 2089d ago Weekly newsletter dedicated to sharing resources for building and operating production machine learning systems. Newsletter #084 2096d ago Weekly newsletter dedicated to sharing resources for building and operating production machine learning systems. Newsletter #083 2103d ago Weekly newsletter dedicated to sharing resources for building and operating production machine learning systems. Newsletter #082 2110d ago Weekly newsletter dedicated to sharing resources for building and operating production machine learning systems.

No matching sources found.