AI Tools & Agents
- OpenAI Agents API
- DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
- Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
- Muse – Meta’s personal AI agent
- Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
- DaVinci Resolve 21.1
- Ask HN: How do you manage skills files?
- GPT-6 Astra on OpenRouter
- Can AI design circuit boards yet?
- Corporate America is getting hooked on open-source AI
- Qwen 3.8 27B available on Cerebras at 1500 tokens/s
- Ask HN: Who is using MCP in production?
- Fable 5.1 World Modeling
- The ChatGPT/Codex app bundles a full copy of LibreOffice
- RAG Is Simpler Than You Think
- Fable and the end of the free lunch
- What Is a Harness?
- I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
- Why your local LLM feels dumber than it is
- New MCP Roadmap
- Munder Difflin – Agent harness to run an office of your clones
- Claudette: Make Claude stop talking like a BuzzFeed article
- DeepSeek-v4-flash-vision-exp
- Show HN: I trained a 125M model to autocomplete piano on-device
- Unsloth Dynamic 3.0 GGUFs
- Qwen 3.8 27B is excellent, but it defaults to overthinking things
- How Organizations Use AI: Evidence from ChatGPT [pdf]
- Accelerating GPT-5.6 Sol Ultrafast
- Mistral OCR 4.1
- Qwen3.8-2.4T
- llama.cpp
- Nvidia Nemotron 3.5 Lightning and NeMo Switchyard
- Grok Bot
- Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
- Humanising LLM Outputs Is Dumb
- Ask HN: What are you working on? (August 2026)
- Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
- Cloudflare OS: an open platform for agents, apps, and work
- DeepSeek V4 Flash on a Single AMD MI300X
- Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
- AI financial advice is surprisingly good, especially if you ask right questions
- qm – Multiplayer agent harness for work
- Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
- The session you cannot take with you
- Agent Skill to Force Docs in ASD-STE100 Simplified Technical English
- Kimi K3-256k
- Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
- Handbook.md shows that long policy documents do not reliably govern agents
- Using an open model feels surprisingly good
- A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
- Claude Cookbook
- Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
- Jack Dorsey launches Buzz to combine team chat, AI agents and Git hosting
- Nativ: Run frontier open models locally on your Mac
- I burned all my tokens researching how to save tokens
- Transcribe.cpp
- Setting up your spare Mac for Claude Code to control, a step-by-step guide
- LM Studio Bionic: the AI agent for open models
- $100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol
- NotebookLM is now Gemini Notebook
- Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU
- Towards a harness that can do anything
- GLM 5.2 is nearly as accurate as a human book keeper
- Show HN: Getting GLM 5.2 running on my slow computer
- GPT‑Live
- We're extending access to Fable 5 on all paid plans through July 12
- Fable turned reMarkable into Tom Riddle's diary from Harry Potter
- Price per 1M tokens is meaningless
- Jamesob's guide to running SOTA LLMs locally
- Fable 5 is Back
- Claude Science
- Computer use in Gemini 3.5 Flash
- For most of the world, open-source AI is the only way forward
- Claude Tag
- Mistral OCR 4
- GLM-5.2 – How to Run Locally
- Sakana Fugu
- There is minimal downside to switching to open models
- Migrate from OpenClaw
- Local Qwen isn't a worse Opus, it's a different tool
- Running local models is good now
- I Fired Google
- Apple Foundation Models
- Ask HN: What are you working on? (June 2026)
- I indexed 669 GB of my GoPro videos using my M1 Max computer and local ML models
- Don't trust large context windows
- RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
- Claude Desktop spawns 1.8 GB Hyper-V VM on every launch, even for chat-only use
- Apache Burr: Build reliable AI agents and applications
- The iPhone's Last Stand?
- Apple Core AI Framework
- Siri AI
- Anthropic, please ship an official Claude Desktop for Linux
- Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency
- Ask HN: What was your "oh shit" moment with GenAI?
- DaVinci Resolve 21
- A 10 year old Xeon is all you need
- 1-Bit Bonsai Image 4B Image Generation for Local Devices
- I put a datacenter GPU in my gaming PC
- Corporate America Is Starting to Ration AI as Cost Skyrockets
- MCP is dead?
- The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin
- Various LLM Smells
- AI sticker shock hits corporate America
- Outsourcing plus local AI will soon become more economical vs. frontier labs
- Uber president says AI spending is getting 'harder to justify'
- DeepSeek makes the V4 Pro price discount permanent
- Show HN: Agent.email – sign up via curl, claim with a human OTP
- Indexing a year of video locally on a 2021 MacBook with Gemma4-31B (50GB swap)
- Qwen3.7-Max: The Agent Frontier
- Gemini 3.5 Flash
- Google I/O
- Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks
- AI is a technology not a product
- Apple Silicon costs more than OpenRouter
- UK sovereign LLM inference
- A few words on DS4
- Claude for Small Business
- Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
- Reimagining the mouse pointer for the AI era
- Running local models on an M4 with 24GB memory
- Ask HN: What are you working on? (May 2026)
- Local AI needs to be the norm
- Using Claude Code: The unreasonable effectiveness of HTML
- Agents need control flow, not more prompts
- DeepSeek 4 Flash local inference engine for Metal
- AlphaEvolve: Gemini-powered coding agent scaling impact across fields
- Higher usage limits for Claude and a compute deal with SpaceX
- Show HN: Tilde.run – Agent sandbox with a transactional, versioned filesystem
- Computer Use is 45x more expensive than structured APIs
- Accelerating Gemma 4: faster inference with multi-token prediction drafters
- Agents for financial services and insurance
- How OpenAI delivers low-latency voice AI at scale
- The agent harness belongs outside the sandbox
- Granite 4.1: IBM's 8B Model Matching 32B MoE
- Claude for Creative Work
- Running local LLMs offline on a ten-hour flight
- The Prompt API
- Show HN: A Karpathy-style LLM wiki your agents maintain (Markdown and Git)
- Google Flow Music
- Could a Claude Code routine watch my finances?
- Website streamed live directly from a model
- Anthropic says OpenClaw-style Claude CLI usage is allowed again
- OpenClaw isn't fooling me. I remember MS-DOS
- Slop Cop
- Claude Design
- Scan your website to see how ready it is for AI agents
- Mozilla Thunderbolt
- The local LLM ecosystem doesn’t need Ollama
- ChatGPT for Excel
- Ask HN: Who is using OpenClaw?
- Google Gemma 4 Runs Natively on iPhone with Full Offline AI Inference
- Turn your best AI prompts into one-click tools in Chrome
- Ask HN: What Are You Working On? (April 2026)
- OpenClaw’s memory is unreliable, and you don’t know when it will break
- I still prefer MCP over skills
- Taste in the age of AI and LLMs
- Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
- Gemma 4 on iPhone
- Tell HN: Anthropic no longer allowing Claude Code subscriptions to use OpenClaw
- April 2026 TLDR Setup for Ollama and Gemma 4 26B on a Mac mini
- Show HN: Apfel – The free AI already on your Mac
- We replaced RAG with a virtual filesystem for our AI documentation assistant
- Lemonade by AMD: a fast and open source local LLM server using GPU and NPU
- Ollama is now powered by MLX on Apple Silicon in preview
- Ensu – Ente’s Local LLM app
- Show HN: Gemini can now natively embed video, so I built sub-second video search
- If DSPy is so great, why isn't anyone using it?
- iPhone 17 Pro Demonstrated Running a 400B LLM
- I built an AI receptionist for a mechanic shop
- Project Nomad – Knowledge That Never Goes Offline
- Flash-MoE: Running a 397B Parameter Model on a Laptop
- MacBook M5 Pro and Qwen3.5 = Local AI Security System
- Nightingale – open-source karaoke app that works with any song on your computer
- Mistral AI Releases Forge
- Kagi Translate now supports LinkedIn Speak as an output language
- Apideck CLI – An AI-agent interface with much lower context consumption than MCP
- My Journey to a reliable and enjoyable locally hosted voice assistant (2025)
- Chrome DevTools MCP (2025)
- The Appalling Stupidity of Spotify's AI DJ
- MCP is dead; long live MCP
- Can I run AI locally?
- Claude now creates interactive charts, diagrams and visualizations
- Show HN: Axe – A 12MB binary that replaces your AI framework
- Personal Computer by Perplexity
- Launch HN: RunAnywhere (YC W26) – Faster AI Inference on Apple Silicon
- Show HN: DenchClaw – Local CRM on Top of OpenClaw
- Show HN: Mcp2cli – One CLI for every API, 96-99% fewer tokens than native MCP
- How to run Qwen 3.5 locally
- Files are the interface humans and agents interact with
- Anthropic, please make a new Slack
- Nvidia PersonaPlex 7B on Apple Silicon: Full-Duplex Speech-to-Speech in Swift
- Glaze by Raycast
- Qwen3.5 Fine-Tuning Guide
- Show HN: I built a sub-500ms latency voice agent from scratch
- OpenClaw surpasses React to become the most-starred software project on GitHub
- WebMCP is available for early preview
- When does MCP make sense vs CLI?
- Why XML tags are so fundamental to Claude
- Switch to Claude without starting over
- Qwen3.5 122B and 35B models offer Sonnet 4.5 performance on local computers
- Show HN: Now I Get It – Translate scientific papers into interactive webpages
- Making MCP cheaper via CLI
- My iPhone 16 Pro Max produces garbage output when running MLX LLMs
- Show HN: LemonSlice – Upgrade your voice agents to real-time video
- MCP is a fad
- Donating the Model Context Protocol and establishing the Agentic AI Foundation
- So you wanna build a local RAG?
- Kagi Assistants
- What if you don't need MCP at all?
- Show HN: Scriber Pro – Offline AI transcription for macOS
- An LLM does not need to understand MCP
- Ollama Turbo
- Show HN: Price Per Token – LLM API Pricing Data
- MCP-B: A Protocol for AI Browser Automation
- Tools: Code Is All You Need
- Launch HN: Issen (YC F24) – Personal AI language tutor
- Harper – an open-source alternative to Grammarly
- Vision Now Available in Llama.cpp
- Suno v4.5
- Claude Integrations
- Claude can now search the web
- MCP vs. API Explained
- Phind 2: AI search with visual answers and multi-step reasoning
- Why LLMs still have problems with OCR
- Show HN: DeepSeek My User Agent
- Qwen2.5-1M: Deploy your own Qwen with context length up to 1M tokens
- Operator research preview
- Hunyuan3D 2.0 – High-Resolution 3D Assets Generation
- Ask HN: Is anyone doing anything cool with tiny language models?
- Show HN: Using YOLO to Detect Office Chairs in 40M Hotel Photos
- Can you read this cursive handwriting? The National Archives wants your help
- Brood War Korean Translations
- Generate audiobooks from E-books with Kokoro-82M
- Has LLM killed traditional NLP?
- GPT-4o with scheduled tasks (jawbone) is available in beta
- LLM based agents as Dungeon Masters
- Adobe Lightroom's AI Remove feature added a Bitcoin to bird in flight photo
- VLC tops 6B downloads, previews AI-generated subtitles
- Agents Are Not Enough
- Apple squandered the Holy Grail
- Show HN: Watch 3 AIs compete in real-time stock trading
- Orbit by Mozilla
- How I run LLMs locally
- Show HN: I made a website to semantically search ArXiv papers
- Building Effective "Agents"
- Tldraw Computer
- The era of open voice assistants
- 1-800-ChatGPT
- Most iPhone owners see little to no value in Apple Intelligence so far
- BlenderGPT
- A ChatGPT clone, in 3000 bytes of C, backed by GPT-2 (2023)
- AI Guesses Your Accent
- Show HN: Cut the crap – remove AI bullshit from websites
- GenChess
- Launch HN: Human Layer (YC F24) – Human-in-the-Loop API for AI Systems
- Show HN: Gemini LLM corrects ASR YouTube transcripts
- Model Context Protocol
- Show HN: FastGraphRAG – Better RAG using good old PageRank
- Kagi Translate
- Project Sid: Many-agent simulations toward AI civilization
- Embeddings are underrated
- Smartphone buyers meh on AI, care more about battery life
- Notes on Anthropic's Computer Use Ability
- Quantized Llama models with increased speed and a reduced memory footprint
- Show HN: Wall-mounted diffusion mirror that turns reflections into paintings
- Show HN: Agent.exe, a cross-platform app to let 3.5 Sonnet control your machine
- Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
- Show HN: HN Update – Hourly news broadcast of top HN stories
- NotebookLM launches feature to customize and guide audio overviews
- Adobe's new image rotation tool is one of the most impressive AI tools seen
- Swarm, a new agent framework by OpenAI
- FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs
- Show HN: A real time AI video agent with under 1 second of latency
- NotebookLM's automatically generated podcasts are surprisingly effective
- Launch HN: Modern Realty (YC S24) – AI Real Estate Agent for Home Buyers
- Forget ChatGPT: why researchers now run small AIs on their laptops
- Scramble: Open-Source Alternative to Grammarly
- Show HN: Bullshit Remover
- Google Illuminate: Books and papers turned into audio
- Show HN: Infinity – Realistic AI characters that can speak
- Phind-405B and faster, high quality AI answers for everyone
- Kagi Assistant
- Web scraping with GPT-4o: powerful but expensive
- 80% of AI Projects Crash and Burn, Billions Wasted Says Rand Report
- Show HN: Remove-bg – open-source remove background using WebGPU
- Anthropic Claude 3.5 can create icalendar files, so I did this
- Midjourney web experience is now open to everyone
- Artificial intelligence is losing hype
- AI companies are pivoting from creating gods to building products
- Launch HN: Trellis (YC W24) – AI-powered workflows for unstructured data
- Show HN: LLM-aided OCR – Correcting Tesseract OCR errors with LLMs
- OTranscribe: A free and open tool for transcribing audio interviews
- Structured Outputs in the API
- Launch HN: Martin (YC S23) – Using LLMs to Make a Better Siri
- Show HN: Turn any website into a knowledge base for LLMs
- OpenAI Announces SearchGPT
- Launch HN: Undermind (YC S24) – AI agent for discovering scientific papers
- When ChatGPT summarises, it does nothing of the kind
- Exo: Run your own AI cluster at home with everyday devices
- Show HN: I generated 70k audiobooks with OpenAI Text-to-Speech
- If AI chatbots are the future, I hate it
- Multi-agent chatbot murder mystery
- Show HN: Tegon: Open-source alternative to Jira, Linear
- Voice Isolator: Strip background noise for film, podcast, interview production
- Meta 3D Gen
- Chrome is adding `window.ai` – a Gemini Nano AI model right inside the browser
- Why we no longer use LangChain for building our AI agents
- Even Apple cannot explain why we need AI in our lives
- McDonald's is ending its drive-thru AI test
- A look at Apple's technical approach to AI including core model performance etc.
- How Alexa dropped the ball on being the top conversational system
- Apple Intelligence for iPhone, iPad, and Mac
- What we've learned from a year of building with LLMs
- Ask HN: How to transcribe 1000s of handwritten notes
- Show HN: ChatGPT UI for rabbit holes
- Vector indexing all of Wikipedia on a laptop
- Ask HN: What is your ChatGPT customization prompt?
- Show HN: We open sourced our entire text-to-SQL product
- Show HN: Route your prompts to the best LLM
- Microsoft Paint's new AI image generator builds on your brushstrokes
- I want flexible queries, not RAG
- Building an AI game studio: what we've learned so far
- Improvements to data analysis in ChatGPT
- Show HN: I made a Mac app to search my images and videos locally with ML
- Claude is now available in Europe
- Player-Driven Emergence in LLM-Driven Game Narrative
- Show HN: I built a non-linear UI for ChatGPT
- Show HN: AI climbing coach – visualize how to climb any route based on your body
- Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU
- An analysis of the Rabbit R1 APK
- Search.chatgpt.com domain and SSL cert have been created
- AI Photo Geolocation
- ESpeak-ng: speech synthesizer with more than one hundred languages and accents
- Show HN: I'm 16 and building an AI based startup called Factful with friends
- Rabbit R1 source code [part 1]
- Claude 3 beats Google Translate
- Show HN: I made a website that converts YT videos into step-by-step guides
- Show HN: Vapi – Convince our voice AI to give you the secret code
- Embeddings are a good starting point for the AI curious app developer
- Ask HN: Is RAG the Future of LLMs?
- Lessons after a Half-billion GPT Tokens
- AI-generated sad girl with piano performs the text of the MIT License
- Udio: Generate music in your favorite styles with a text prompt
- Show HN: Sonauto – A more controllable AI music creator
- World_sim: LLM prompted to act as a sentient CLI universe simulator
- AI-generated sad girl with piano performs the text of the MIT License
- Show HN: I've built a locally running Perplexity clone
- LLaMA now goes faster on CPUs
- Ask HN: Is anybody getting value from AI Agents? How so?
- OpenVoice: Versatile instant voice cloning
- Launch HN: Aqua Voice (YC W24) – Voice-driven text editor
- OpenAI GPT-4 vs. Groq Mistral-8x7B
- Launch HN: Soundry AI (YC W24) – Music sample generator for music creators
- Ask HN: If you've used GPT-4-Turbo and Claude Opus, which do you prefer?
- Ollama now supports AMD graphics cards
- TextSnatcher: Copy text from images, for the Linux Desktop
- Show HN: Skyvern – Browser automation using LLMs and computer vision
- Spreadsheets are all you need
- A generalist AI agent for 3D virtual environments
- Show HN: I made an app to use local AI as daily driver
- Show HN: AI dub tool I made to watch foreign language videos with my 7-year-old
- Gemma.cpp: lightweight, standalone C++ inference engine for Gemma models
- Brave's AI assistant now integrates with PDFs and Google Drive
- Show HN: Real-time image generation with SDXL Lightning
- Launch HN: Danswer (YC W24) – Open-source AI search and chat over private data
- Launch HN: Retell AI (YC W24) – Conversational Speech API for Your LLM
- Groq runs Mixtral 8x7B-32k with 500 T/s
- Ollama is now available on Windows in preview
- Ask HN: What are some actual use cases of AI Agents right now?
- Show HN: Reor – An AI note-taking app that runs models locally
- Memory and new controls for ChatGPT
- Nvidia's Chat with RTX is an AI chatbot that runs locally on your PC
- OpenAI compatibility
- Ask HN: What have you built with LLMs?
- Show HN: Natural-SQL-7B, a strong text-to-SQL model
- A new way to discover places with generative AI in Maps
- Show HN: A simple ChatGPT prompt builder
- Launch HN: Univerbal (YC W23) – Language learning with a conversational AI tutor
- Show HN: Phrasing – learn every language, to any level
- Show HN: WhisperFusion – Low-latency conversations with an AI chatbot
- Brave Leo now uses Mixtral 8x7B as default
- Ollama releases Python and JavaScript Libraries
- My AI costs went from $100 to less than $1/day: Fine-tuning Mixtral with GPT4
- WhisperSpeech – An open source text-to-speech system built by inverting Whisper
- Field experimental evidence of AI on knowledge worker productivity and quality
- Vanna.ai: Chat with your SQL database
- Building a fully local LLM voice assistant to control my smart home
- Changes we're making to Google Assistant
- The GPT Store
- Rabbit: LLM-First Mobile Phone
- Show HN: Auto Wiki – Turn your codebase into a Wiki
- I made an app that runs Mistral 7B 0.2 LLM locally on iPhone Pros
- Launch HN: Rosebud (YC S19) – Turn game descriptions into browser games
- Show HN: Inbox Zero – open-source email assistant
- Show HN: Rem: Remember Everything (open source)
- Ask HN: How do I train a custom LLM/ChatGPT on my own documents in Dec 2023?
- Suno AI
- Groqchat
- Labs.Google
- Apple wants AI to run directly on its hardware instead of in the cloud
- Mistral 7B Fine-Tune Optimized
- Perplexity Labs Playground
- Google's True Moonshot
- Prompt engineering
- First look at Microsoft 365 Copilot
- Bash one-liners for LLMs
- Google Gemini Pro API Available Through AI Studio
- MemoryCache: Augmenting local AI with browser data
- Role-playing with AI will be a powerful tool for writers and educators
- Show HN: Open-source macOS AI copilot using vision and voice
- Mistral: Our first AI endpoints are available in early access
- Show HN: I Remade the Fake Google Gemini Demo, Except Using GPT-4 and It's Real
- Show HN: A Dalle-3 and GPT4-Vision feedback loop
- LM Studio – Discover, download, and run local LLMs
- ChatGPT with voice is now available to all free users
- Krita AI Diffusion
- StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
- Frigate: Open-source network video recorder with real-time AI object detection
- Death by AI – a free Jackbox style party game. AI judges your plans to survive
- Exploring GPTs: ChatGPT in a trench coat?
- Ask HN: Is anyone else bearish on OpenAI?
- Real-time image editing using latent consistency models
- Don't build AI products the way everyone else is doing it
- Using GPT-4 Vision with Vimium to browse the web
- AI can catalogue a forest's inhabitants simply by listening