Models & Research
- I built non-autoregressive decision models with RL a year ago
- GPT-6 Astra Solves a WWI German Radio Cipher
- DeepSeek v4.1 Flash
- GPT-6 Astra, looped transformers, and hidden reasoning
- DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
- ChatGPT Images 2.5
- AlphaGenome Atlas: a high-resolution map of human DNA
- Nvidia's Jensen Huang says 'AGI has arrived' and congratulates OpenAI
- “Next-token predictor” is the wrong mental model for LLMs
- OpenAI's GPT-6 Astra on ARC-AGI-3
- GPT-6 Astra
- OpenAI begins rolling out GPT-6 Astra
- How concerned should we be about Astra's recurrent architecture?
- K2 Horizon: A connected fleet of six open models
- Go grandmaster Shin defeats AI KataGo with a two-stone handicap
- Muse Spark 1.3
- Quasar 438B: Europe's Leading AI Model
- The Emergent Symbolic Structure of Artificial Neural Networks
- Claude Fable 5.1 and Claude Mythos 5.1
- Qwen3.8-Flash-Next
- Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
- Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
- Fable and the end of the free lunch
- GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
- Ox Alpha
- If this is true, the hyperscalers are toast
- Opus 5.0 drives incoherence into the stratosphere
- Qwen3.8 27B scores 52 on Artificial Analysis
- GPT 5.6 Sol is the best "vision" model OpenAI ever released
- Models Are Getting Dumber on Purpose
- What happens when an LLM never sees material beyond fifth grade?
- Gemini 3.7 Flash
- Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
- Grok 4.6
- What sort of maths are LLMs good at?
- Compression is prediction
- Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models
- Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
- DeepMind's WeatherNext model achieves breakthrough forecasting cyclones
- U.S. Department of Energy Launches the Genesis Open Models Initiative
- Qwen3.8 Max now ranked as the best overall model by agentic index
- Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users
- Position: LLMs Can't Jump
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
- Karpathy’s Pelican
- Seedance 2.5
- DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
- DeepSeek-V4-Flash Update
- Kimi K3-256k
- Some thoughts about Anthropic's new cryptanalysis results
- A walk through of the DeltaNet family of linear attention variants
- Kimi K3 Architecture Overview and Notes
- Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
- A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
- Kimi-K3 Technical Report [pdf]
- Kimi-K3 on HuggingFace
- ARC-AGI Leaderboard
- Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
- Flux 3
- Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
- "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
- Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge
- Xiaomi-Robotics-1
- Moonshot AI suspends new subscriptions due to Kimi K3 demand
- Qwen 3.8
- The Kimi K3 Moment
- The state of open source AI
- Kimi K3, and what we can still learn from the pelican benchmark
- Kimi K3: Open Frontier Intelligence
- Inkling: Our Open-Weights Model
- Bonsai 27B: A 27B-Class model that runs on a phone
- How to stop Claude from saying load-bearing
- Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor
- Hy3
- Muse Spark 1.1
- Mistral's Robostral Navigate: a state of the art robotics navigation model
- GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday
- A global workspace in language models
- What Emily Bender meant by "stochastic parrots"
- Leanstral 1.5
- Claude Sonnet 5
- Nano Banana 2 Lite
- GLM 5.2 beats Claude in our benchmarks
- Asian AI startups launch Mythos-like models
- DSpark: Speculative decoding accelerates LLM inference [pdf]
- The gap between open weights LLMs and closed source LLMs
- Previewing GPT‑5.6 Sol: a next-generation model
- Unlimited OCR: One-shot long-horizon parsing
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
- Apertus – Open Foundation Model for Sovereign AI
- There is minimal downside to switching to open models
- GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
- DeepSeek Introduces Vision
- GLM-5.2 is the new leading open weights model on Artificial Analysis
- GPT‑NL: a sovereign language model for the Netherlands
- Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
- GLM 5.2 Is Out
- Rich Sutton on AI creativity and discovery
- Claude Fable 5
- MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
- If LLMs Have Human-Like Attributes, Then So Does Age of Empires II
- They’re made out of weights
- Gemma 4 12B: A unified, encoder-free multimodal model
- Claude Opus 4.8
- Disagreement among frontier LLMs on real-world fact-checks
- All of human cooking compressed into 2 megabytes
- A sleep-like consolidation mechanism for LLMs
- Norway's 2 petabytes of Huawei flash storage and LLM training
- Gemini Omni
- Show HN: Gaussian Splat of a Strawberry
- GenCAD
- SANA-WM, a 2.6B open-source world model for 1-minute 720p video
- The sigmoids won't save you
- Natural Language Autoencoders: Turning Claude's Thoughts into Text
- Accelerating Gemma 4: faster inference with multi-token prediction drafters
- Grok 4.3
- Granite 4.1: IBM's 8B Model Matching 32B MoE
- Mistral Medium 3.5
- VibeVoice: Open-source frontier voice AI
- Talkie: a 13B vintage language model from 1930
- OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
- There Will Be a Scientific Theory of Deep Learning
- DeepSeek v4
- GPT-5.5
- ChatGPT Images 2.0
- Kimi K2.6: Advancing open-source coding
- Claude Opus 4.7
- Muse Spark: Scaling towards personal superintelligence
- GLM-5.1: Towards Long-Horizon Tasks
- Show HN: I built a tiny LLM to demystify how language models work
- Google releases Gemma 4 open models
- The case for zero-error horizons in trustworthy LLMs
- Qwen3.6-Plus: Towards real world agents
- Show HN: 1-Bit Bonsai, the First Commercially Viable 1-Bit LLMs
- Google's 200M-parameter time-series foundation model with 16k context
- ARC-AGI-3
- TurboQuant: Redefining AI efficiency with extreme compression
- Epoch confirms GPT5.4 Pro solved a frontier math open problem
- Show HN: Three new Kitten TTS models – smallest less than 25MB
- Measuring progress toward AGI: A cognitive framework
- Why AI systems don't learn – On autonomous learning from cognitive science
- GPT‑5.4 Mini and Nano
- Executing programs inside transformers with exponentially faster inference
- BitNet: Inference framework for 1-bit LLMs
- Show HN: How I topped the HuggingFace open LLM leaderboard on two gaming GPUs
- Yann LeCun raises $1B to build AI that understands the physical world
- GPT-5.4
- GPT‑5.3 Instant
- Why XML tags are so fundamental to Claude
- Microgpt
- Nano Banana Pro
- A new Google model is nearly perfect on automated handwriting recognition
- DeepSeek OCR
- Which table format do LLMs understand best?
- Markov chains are the original language models
- Releasing weights for FLUX.1 Krea
- Complete silence is always hallucinated as "ترجمة نانسي قنقر" in Arabic
- Seven replies to the viral Apple reasoning paper and why they fall short
- Chatterbox TTS
- Vision Language Models Are Biased
- Suno v4.5
- GPT-4.1 in the API
- Why LLMs still have problems with OCR
- Run DeepSeek R1 Dynamic 1.58-bit
- Open-R1: an open reproduction of DeepSeek-R1
- The Illustrated DeepSeek-R1
- DeepSeek releases Janus Pro, a text-to-image generator [pdf]
- Explainer: What's r1 and everything else?
- Emerging reasoning with reinforcement learning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
- Results of "Humanity's Last Exam" benchmark published
- OpenAI's o1 Playing Codenames
- Tensor Product Attention Is All You Need
- Hunyuan3D 2.0 – High-Resolution 3D Assets Generation
- DeepSeek-R1
- FrontierMath was funded by OpenAI
- Minecraft with object impermanence
- O1 isn't a chat model (and that's the point)
- Phi 4 available on Ollama
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English? (2023)
- 30% drop in O1-preview accuracy when Putnam problems are slightly variated
- Deepseek: The quiet giant leading China’s AI race
- Coconut by Meta AI – Better LLM Reasoning with Chain of Continuous Thought?
- Does current AI represent a dead end?
- Can AI do maths yet? Thoughts from a mathematician
- GPT-5 is behind schedule
- OpenAI O3 breakthrough high score on ARC-AGI-PUB
- Cultural Evolution of Cooperation Among LLM Agents
- Veo 2: Our video generation model
- Ilya Sutskever NeurIPS talk [video]
- New LLM optimization technique slashes memory costs
- Phi-4: Microsoft's Newest Small Language Model Specializing in Complex Reasoning
- A ChatGPT clone, in 3000 bytes of C, backed by GPT-2 (2023)
- Gemini 2.0: our new AI model for the agentic era
- Training LLMs to Reason in a Continuous Latent Space
- Sora is here
- Google says AI weather model masters 15-day forecast
- Llama-3.3-70B-Instruct
- Genie 2: A large-scale foundation world model
- Amazon Nova
- World Labs: Generate 3D worlds from a single image
- Procedural knowledge in pretraining drives reasoning in large language models
- The Curse of Recursion: Training on generated data makes models forget (2023)
- QwQ: Alibaba's O1-like reasoning LLM
- DeepThought-8B: A small, capable reasoning model
- What happens if we remove 50 percent of Llama?
- OK, I can partly explain the LLM chess weirdness now
- Extending the context length to 1M tokens
- Something weird is happening with LLMs and chess
- Francois Chollet is leaving Google
- OpenAI, Google and Anthropic are struggling to build more advanced AI
- LLMs have reached a point of diminishing returns
- FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI
- Perceptually lossless (talking head) video compression at 22kbit/s
- The deep learning boom caught almost everyone by surprise
- Tencent Hunyuan-Large
- Chain-of-thought can hurt performance on tasks where thinking makes humans worse
- LLMs know more than they show: On the intrinsic representation of hallucinations
- A return to hand-written notes by learning to read and write
- Detecting when LLMs are uncertain
- AI engineers claim new algorithm reduces AI power consumption by 95%
- Grandmaster-level chess without search
- Use Prolog to improve LLM's reasoning
- Diffusion for World Modeling
- FLUX is fast and it's open source
- LLMs don't do formal reasoning
- Understanding the Limitations of Mathematical Reasoning in LLMs
- Addition is all you need for energy-efficient language models
- Differential Transformer
- LLMs, Theory of Mind, and Cheryl's Birthday
- Meta Movie Gen
- The real data wall is billions of years of evolution
- Were RNNs all we needed?
- FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs
- Liquid Foundation Models: Our First Series of Generative AI Models
- AGI is far from inevitable
- If AI seems smarter, it's thanks to smarter human trainers
- Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
- Two new Gemini models, reduced 1.5 Pro pricing, increased rate limits, and more
- Why wordfreq will not be updated
- Chain of Thought empowers transformers to solve inherently serial problems
- Launch HN: Silurian (YC S24) – Simulate the Earth
- g1: Using Llama-3.1 70B on Groq to create o1-like reasoning chains
- Terence Tao on O1
- OpenAI o1 Results on ARC-AGI-Pub
- Notes on OpenAI's new o1 chain-of-thought models
- Learning to Reason with LLMs
- GPTs and Hallucination
- Radiology-specific foundation model
- How does cosine similarity work?
- Inductive or deductive? Rethinking the fundamental reasoning abilities of LLMs
- Update on Llama adoption
- Diffusion models are real-time game engines
- Anthropic publishes the 'system prompts' that make Claude tick
- Transformers in music recommendation
- Markov chains are funnier than LLMs
- Are you better than a language model at predicting the next word?
- Does Reasoning Emerge? Probabilities of Causation in Large Language Models
- Grok-2 Beta Release
- RLHF is just barely RL
- Flux: Open-source text-to-image model with 12B parameters
- Calculating the cost of a Google DeepMind paper
- SAM 2: Segment Anything in Images and Videos
- AI solves International Math Olympiad problems at silver medal level
- AI models collapse when trained on recursively generated data
- Large Enough
- Open source AI is the path forward
- Llama 3.1
- Mistral NeMo
- Want to spot a deepfake? Look for the stars in their eyes
- Overcoming the limits of current LLMs
- Large models of what? Mistaking engineering achievements for linguistic agency
- Vision language models are blind
- Reasoning in Large Language Models: A Geometric Perspective
- Tokens are a big reason today's generative AI falls short
- A Model of a Mind
- Rodney Brooks on limitations of generative AI
- Gemma 2: Improving Open Language Models at a Practical Size [pdf]
- Open-Sora does pretty good video generation on consumer GPUs
- Why your brain is 3 milion more times efficient than GPT-4
- Claude 3.5 Sonnet
- Getting 50% (SoTA) on Arc-AGI with GPT-4o
- AI Search: The Bitter-Er Lesson
- MLow: Meta's low bitrate audio codec
- ARC Prize – a $1M+ competition towards open AGI progress
- How Does GPT-4o Encode Images?
- Extracting concepts from GPT-4
- Qwen2 LLM Released
- Stable Audio Open
- Simple tasks showing reasoning breakdown in state-of-the-art LLMs
- LLMs aren't "trained on the internet" anymore
- Re-Evaluating GPT-4's Bar Exam Performance
- “Imprecise” language models are smaller, speedier, and nearly as accurate
- Training is not the same as chatting: LLMs don’t remember everything you say
- Reproducing GPT-2 in llm.c
- Transformers Can Do Arithmetic with the Right Embeddings
- Financial Statement Analysis with Large Language Models
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
- Llama3 implemented from scratch
- I'm Bearish OpenAI
- Toon3D: Seeing cartoons from a new perspective
- ChatGPT-4o vs. Math
- PaliGemma: Open-Source Multimodal Model by Google
- Gemini Flash
- Veo
- GPT-4o's Memory Breakthrough – Needle in a Needlestack
- GPT-4o
- Falcon 2
- No "Zero-Shot" Without Exponential Data
- AlphaFold 3 predicts the structure and interactions of life's molecules
- TimesFM: Time Series Foundation Model for time-series forecasting
- LLMs can't do probability
- Better and Faster Large Language Models via Multi-Token Prediction
- Kolmogorov-Arnold Networks
- GPT-4.5 or GPT-5 being tested on LMSYS?
- What can LLMs never do?
- Snowflake Arctic Instruct (128x3B MoE), largest open source model
- The question that no LLM can answer and why it is important
- CoreNet: A library for training deep neural networks
- VideoGigaGAN: Towards detail-rich video super-resolution
- Phi-3 Technical Report
- Llama 3 8B is almost as good as Wizard 2 8x22B
- Meta Llama 3
- Cyc: History's Forgotten AI Project
- Show HN: Speeding up LLM inference 2x times (possibly)
- Mixtral 8x22B
- Google DeepMind's Aloha Unleashed is pushing the boundaries of robot dexterity
- Visualizing Attention, a Transformer's Heart [video]
- Scaling will never get us to AGI
- Show HN: Next-token prediction in JavaScript
- Mistral AI Launches New 8x22B MOE Model
- GPT-4 Turbo with Vision Generally Available
- Llm.c – LLM training in simple, pure C/CUDA
- More Agents Is All You Need: LLMs performance scales with the number of agents
- Choose your weapon: Survival strategies for depressed AI academics
- LLMs use a surprisingly simple mechanism to retrieve some stored knowledge
- Towards 1-bit Machine Learning Models
- DBRX: A new open LLM
- Is GPT-4 a good data analyst? (2023)
- “Emergent” abilities in LLMs actually develop gradually and predictably – study
- GPT-4, without specialized training, beat a GPT-3.5 class model that cost $10M
- How Chain-of-Thought Reasoning Helps Neural Networks Compute
- The Google employees who created transformers
- Stability.ai – Introducing Stable Video 3D
- How do neural networks learn?
- Grok
- Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
- Large Language Models Are Neurosymbolic Reasoners
- Is Cosine-Similarity of Embeddings Really About Similarity?
- OpenAI – transformer debugger release
- How far are we from intelligent visual deductive reasoning?
- Fine tune a 70B language model at home
- What if AGI is not coming?
- Training LLMs from ground zero as a startup
- Why do tree-based models still outperform deep learning on tabular data? (2022)
- Opus 1.5 released: Opus gets a machine learning upgrade
- Claude 3 model family
- Top AIs still fail IQ tests
- The Era of 1-bit LLMs: ternary parameters for cost-effective computing
- Mistral Large
- Hallucination is inevitable: An innate limitation of large language models
- Every model learned by gradient descent is approximately a kernel machine (2020)
- Generative Models: What do they know? Do they know things? Let's find out
- I Spent a Week with Gemini Pro 1.5–It's Fantastic
- Beyond A*: Better Planning with Transformers
- Stable Diffusion 3
- The killer app of Gemini Pro 1.5 is using video as an input
- Gemma: New Open Models
- Jeff Dean: Trends in Machine Learning [video]
- Let me clear a huge misunderstanding
- Magika: AI powered fast and efficient file type identification
- Video generation models as world simulators
- Sora: Creating video from text
- Our next-generation model: Gemini 1.5
- Stable Cascade
- Stable-Audio-Demo
- Tiny quadrotor learns to fly in 18 seconds
- Grandmaster-Level Chess Without Search
- Apple releases MGIE, an AI-based image editing model
- Bard's latest updates: Access Gemini Pro globally and generate images
- Markov Chains are the Original Language Models
- Mistral CEO confirms 'leak' of new open source AI model nearing GPT4 performance
- Lumiere: A space-time diffusion model for realistic video generation
- New theory suggests LLMs can understand text
- GPT-3.5 crashes when it thinks about useRalativeImagePath too much
- If you can't reproduce the model then it's not open-source
- AlphaGeometry: An Olympiad-level AI system for geometry
- Mixtral 8x7B: A sparse Mixture of Experts language model
- Chess-GPT's Internal World Model
- Understand how transformers work by demystifying the math behind them
- Stuff we figured out about AI in 2023
- Will scaling work?
- How many legs do ten elephants have, if two of them are legless?
- Ferret: A Multimodal Large Language Model
- Mistral 7B Fine-Tune Optimized
- Implementation of Mamba in one file of PyTorch
- Advancements in machine learning for machine learning
- FunSearch: Making new discoveries in mathematical sciences using LLMs
- SMERF: Streamable Memory Efficient Radiance Fields
- Google Imagen 2
- Artificial intelligence systems found to excel at imitation, but not innovation
- Phi-2: The surprising power of small language models
- Mixtral of experts
- Show HN: I Remade the Fake Google Gemini Demo, Except Using GPT-4 and It's Real
- Three things that LLMs have made us rethink
- Mistral "Mixtral" 8x7B 32k model [magnet]
- Let's try to understand AI monosemanticity
- $10M AI Mathematical Olympiad Prize
- Please ignore the deluge of complete nonsense about Q*
- MonadGPT – What would have happened if ChatGPT was invented in the 17th century?
- Training for one trillion parameter model backed by Intel and US govt has begun
- AI system self-organises to develop features of brains of complex organisms
- About That OpenAI "Breakthrough"
- OpenAI researchers warned board of AI breakthrough ahead of CEO ouster
- Stable Video Diffusion
- Claude 2.1
- Exponentially faster language modelling
- LLMs cannot find reasoning errors, but can correct them
- Comparing humans, GPT-4, and GPT-4V on abstraction and reasoning tasks
- I disagree with Geoff Hinton regarding "glorified autocomplete"
- LLMs by Hallucination Rate
- Is the reversal curse in LLMs real?
- GraphCast: AI model for weather forecasting