GLM-5.3 发布:基座不变,编程能力却大幅跃升
智谱今日发布 GLM-5.3,与前代不同,这次基座模型未变,所有提升都来自后训练阶段的扩展。通过十倍规模的长任务环境、更多样的环境类型和更长的训练时间,模型智能上限被显著拉高。内部评测显示,编程体验比 GLM-5.2 提升约 50%,成为目前编程能力最强的开源模型。公开基准测试中,Terminal-Bench 3.0 得分从 4.6 飙升至 28.3,DeepSWE v1.1 从 46.2 升至 66.9,表现逼近闭源旗舰。

Discover
Language
Account
Daily discover the most amazing AI world - from breakthrough news to innovative products, from cutting-edge projects to tech trends
智谱今日发布 GLM-5.3,与前代不同,这次基座模型未变,所有提升都来自后训练阶段的扩展。通过十倍规模的长任务环境、更多样的环境类型和更长的训练时间,模型智能上限被显著拉高。内部评测显示,编程体验比 GLM-5.2 提升约 50%,成为目前编程能力最强的开源模型。公开基准测试中,Terminal-Bench 3.0 得分从 4.6 飙升至 28.3,DeepSWE v1.1 从 46.2 升至 66.9,表现逼近闭源旗舰。

苹果在中国市场的AI战略迎来重大调整。据内部消息,苹果已与阿里巴巴合作,专门为中国市场训练了一款大语言模型。这一举措打破了苹果以往依赖海外模型的惯例,也使其成为首家在中国获得监管批准并推出专属大模型的外国公司。苹果的AI工具套件Apple Intelligence预计将在数月后随iOS更新正式登陆中国。
腾讯大模型布局又有新动作。据“智能涌现”报道,腾讯元老级研究员许灿已正式调任微信事业群,加入WeLM大模型团队,专注于后训练与Agent开发。许灿是微软亚研院出身,曾主导WizardLM项目并提出Evol-Instruct方法,在合成数据、强化学习反馈等领域经验丰富。此次调整正值微信AI加速落地之际,其AI助手“小微”已开始小范围灰度测试,WeLM团队也在推进稀疏MoE架构,以降低推理成本。许灿的加入有望增强微信AI的任务规划与工具调用能力,为微信在C端AI入口竞争中开辟新路径。
美国司法部在谷歌反垄断案中提出新动议,试图禁止谷歌向浏览器支付默认搜索引擎费用。这一举措本意是打破谷歌的垄断,却意外地将Firefox浏览器母公司Mozilla推入生存困境。Mozilla约85%的美国市场收入依赖谷歌的这笔付款,一旦失去,其业务将难以为继。这场法律博弈充满了讽刺意味,独立浏览器厂商的未来悬于一线。
Jia Zhangke's new film 'Dunhuang Mama' has been registered, telling the story of a lonely mother in Dunhuang who finds solace in AI and embarks on a cross-country journey. The film, starring Zhao Tao and Liao Fan, is expected in 2027. It marks a significant move of AI into serious auteur cinema, exploring technology's role in connecting rather than isolating.
Zhipu AI has released GLM-5.3, a 740-billion-parameter model that keeps its size but boosts performance by 50% through post-training. It now rivals Claude's coding abilities and tops open-source benchmarks. Available on Zhipu's tools and third-party platforms, with open-source weights coming soon.

At Baidu's AI Day, the company unveiled the Chinese name for its general-purpose AI agent, GenFlow, now called Kuku AI. With over 100 million monthly active users, Baidu is spinning off Kuku AI into a standalone product suite, including PC, web, mini program, and enterprise editions. This move signals a stronger push into AI-powered office tools, aiming to cover everything from personal productivity to corporate collaboration.
Kingsoft Office's AI assistant, Lingxi Professional Edition, has integrated DeepSeek's latest V4-Pro model, unlocking a million-token context window. This means the AI can now handle complex office work by understanding entire document histories and user preferences. The update promises more stable and autonomous task execution, delivering editable, traceable results directly in Office apps. Available now for all users, this marks a significant step toward AI that truly collaborates on professional projects.
京东旗下七鲜咖啡宣布,全球首家24小时智能无人咖啡店将于8月16日在北京银河SOHO开业。店内从萃取到出品全程由机器人和AI完成,顾客可自由调节咖啡浓度、甜度等,还能为饮品命名。开业当天将进行24小时机器人制作直播,并挑战吉尼斯世界纪录。此外,AI个性化推荐功能即将上线,可根据用户身体状况建议咖啡因摄入量。这一模式能否持续吸引复购,仍有待市场检验。

本地部署的模型真的能应对复杂业务吗?蚂蚁数科团队与开源社区合作,通过AReno工具包和百灵大模型,在单机环境下成功建立了智能体强化学习闭环。他们用井字棋作为验证任务,让模型在400步训练后实现了质的飞跃。这一方法不仅适用于棋类游戏,还能迁移到工具调用修复、结构化字段提取等真实场景。

DeepSeek's latest model, V4 Pro, just landed on SiliconFlow with a massive 1M token context window and three reasoning intensity levels. But the real headline? Cache-hit tokens cost just $0.44 per million—a game-changer for developers running agent workflows. This launch keeps the MIT license, making it a tempting option for teams watching their budgets. Here's what you need to know about pricing, performance, and why this matters for long-context tasks.
After nearly six years, Apple's HomePod mini is finally getting a sequel. The second-generation model is expected to launch this autumn, featuring a faster chip to support the new AI-powered Siri. This update isn't just about hardware—it's a test of Apple's edge AI capabilities. With better UWB connectivity and improved audio, the HomePod mini 2 aims to shake off its reputation as the weakest voice assistant. Will it succeed? We'll find out soon.
8月14日,一汽-大众全新纯电车型ID. AURA T6开启盲订,并确认接入豆包大模型。这标志着双方在智能座舱领域的合作从燃油车扩展至纯电领域,为智能出行带来更多想象空间。
Launched on August 10, the Qwen Open Platform has already drawn over 200 ecosystem partners, including major names like SF Express and KE Holdings. The platform simplifies AI agent integration, allowing partners to offer services across logistics, real estate, and more. With built-in tools and one-click authorization, it's lowering barriers and accelerating the shift from standalone tools to direct service access points.

MiniMax has unveiled Music3, a music generation model that turns lyrics and descriptions into full-length songs up to five minutes long. Using a clever two-part AI system, it handles everything from the song's structure to its finest audio details. The result? A 32kHz stereo WAV file that could be your next favorite tune. Here's how it works and why it matters for music lovers and creators alike.

微软宣布,Copilot语音模式中的黄色小怪物Mico将退出聊天界面,转至Learn Live学习平台。这一变动距其发布不到一年,背后是Copilot与Microsoft 365 Copilot的合并计划。Mico的离开,再次让人想起Clippy、Cortana等被微软放弃的虚拟助手。不过,Mico并非一无所获,它验证了语音交互中“拟人化表达”的价值,为Copilot的未来形态留下了参考。
Guangdong has introduced a new financial product called 'Token Loan' to help AI and computing power companies secure financing. The loan amount is based on a company's token consumption and contracts, offering a fresh way for startups to get credit. This marks a significant step in integrating token economy into traditional banking.
Google has unveiled Gemini 3.7 Flash, an upgraded AI model that promises to transform coding and app development. With superior code debugging and production-ready output, it cuts token costs and boosts efficiency. Starting today, it powers Gemini Spark, making developers' lives easier. Here's what you need to know.
OpenAI has finally rolled out the 'Ultra-Fast' mode for GPT-5.6 Sol, promising a 14-fold speed boost. But this isn't just about raw speed—it's a clever hardware play with Cerebras. The catch? Access is limited, and only certain customers will get it. Here's what you need to know.

A developer has exposed a troubling trend on the App Store: apps using AI-generated fake screenshots to display inflated ratings, like a 4.9-star score, are slipping through Apple's review process. This raises serious questions about the effectiveness of Apple's quality checks and the potential privacy risks, especially for Safari extensions that can access user data.
