Machine Learning Research
发布时间:2026-09-08 | 浏览:1
A Simpler Method to Monitor Models: CRC Monitor can tell if model output is incorrect or unsafe in real-time
LLM safety monitors that act during generation often analyze a series of safety scores to catch incorrect or harmful outputs.
Custom Models for Law, News, and Finance: Thomson Reuters’ Thomson LLM, based on Qwen 3.5, trained on proprietary data
Thomson Reuters launched a proprietary large language model family called Thomson.
Ox Alpha Revealed as GLM-5.3-Flash: Z.ai’s latest model built viral buzz as Ox Alpha while being served on China-made chips
For over a week, the name and maker of the most-used model on OpenRouter remained unknown.
LLMs Take Out the Agent’s Trash: Researchers at Xiaohongshu detail LLM-based memory management technique for agents
AI agents typically compact the contents of their context windows by summarizing or deleting the oldest material.
DeepSeek-V4-Pro Gets Refreshed: DeepSeek’s flagship only slightly outperforms Flash model; Open source harness draws interest
DeepSeek’s flagship model graduated from preview with improved performance plus the harness the model was benchmarked in.
Inside AI’s Need for Speed: OpenAI Partners with Cerebras, Google Releases Gemini 3.7 Flash, Nvidia Debuts Nemotron 3.5 Lightning
Developers who already track models’ cost and accuracy have good new reasons to pay closer attention to a third essential factor: speed.
GLM-5.3 Makes Cybersecurity Gains: Z.ai delayed weights for GLM-5.3 due to cybersecurity risk
Z.ai’s latest flagship model effectively ties open-weights leader Kimi K3 on Artificial Analysis’ index of intelligence benchmarks.
Agents Come to Speech Recognition: AgenticASR incorporates user corrections to edit speech recognition on the fly
Most speech-to-text systems transcribe speech in a single pass, which makes them unable to correct errors in their outputs. Researchers built a system that allows for interactive corrections.
Qwen3.8-Max Lands With A Bang: Inside Alibaba's two new models, Qwen3.8-Max and Qwen3.8-27B
Open models are getting larger and more capable. A few weeks ago, Moonshot AI announced Kimi K3, the largest and best-performing open weights model yet. Last week, Alibaba answered by releasing weights for a giant of its own.
How Claude’s Watermarks Work: Anthropic details new invisible marking system for future versions of Claude
Anthropic introduced invisible, machine-readable signals that text and images were generated by Claude.
Grok’s Cursor Alliance Pays Off: Grok 4.6 rivals Claude Opus 5 and GPT-5.6 Sol at a lower price
Once a lab that produced mid-tier models, SpaceXAI has steadily improved. It just built one of the most capable models in the world while keeping prices relatively low.
AI Can Help Heal Romantic Distress: Inside overit, a study of chatbot therapy after breakups
Chatbots that are designed to treat mental health issues typically require multiple sessions, posing a risk that users will drop out before they receive much benefit. Researchers showed that chatbots can provide relief in a single session.
Minimax’s State-of-the-Art Video Model Is Only Minimally Open: Minimax H3's weights are free, but carry unusual restrictions
A free-to-download model sets a new standard for video generation and editing, but its license comes with unexpected restrictions.
Google’s Robotics Model Has Legs: Gemini Robotics 2 gets Google closer to true multi-embodiment
Google’s latest vision-language-action model can walk a humanoid robot across a room, crouch to a low shelf, and close a five-fingered hand around a light bulb.
Muse Code Wants Your Data: Meta's Muse Spark 1.2 and Muse Code approach the intelligence frontier at a discount
Meta will cut coding bills from dollars to pennies for developers who let the company learn from their work.
Subscribe to The Batch
Stay updated with weekly AI News and Insights delivered to your inbox
Choose Your Plan