X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(519)
@op7418@op7418AI 评分5858 
@berryxia@berryxiaAI 评分4040
引用Daily Dose of Data Science (@DailyDoseOfDS_)@DailyDoseOfDS_Transformer and Mixture of Experts, explained visually! Mixture of Experts (MoE) is a popular architecture that uses different experts to improve Transformer models. Transformer and MoE differ in the decoder block: - Transformer uses a feed-forward network. - MoE uses experts, which are feed-forward networks but smaller compared to those Transformer. During inference, a subset of experts are selected. This makes inference faster in MoE. Also, since the network has multiple decoder layers: - The text passes through different experts across layers. - The chosen experts also differ between tokens. But how does the model decide which experts should be ideal? The router does that. It is a multi-class classifier that produces softmax scores over experts to select the top K experts. The router is trained with the network, and it learns to select the best experts. But it isn't straightforward. There are challenges! Challenge 1) Notice this pattern at the start of training: - Say, the model selects "Expert 2" - This expert gets a bit better - It may get selected again since it's the "best" - It learns more - It gets selected again in the next iteration - It learns more, and so on! This means many experts can go under-trained due to the overselection of a few experts! We solve this in two steps: - Add noise to the feed-forward output of the router so that other experts can get higher logits. - Set all but the top K logits to -infinity. After softmax, these scores become zero. This way, other experts also get the opportunity to train. Challenge 2) Some experts may get exposed to more tokens than others, leading to under-trained experts. We prevent this by limiting the number of tokens an expert can process. If an expert reaches the limit, the token is passed to the next best expert. Overall, MoEs have more parameters to load. But a fraction of them are activated during inference. This leads to faster inference. Mixtral 8x7B and Llama 4 are two popular MoE-based LLMs. Have you used MoEs in production yet?
@kimmonismus@kimmonismusAI 评分6161 美国 10 年期国债收益率升至 4.568%,为 10 个月来最高,30 年期重回 5% 以上;通胀重新加速,市场已完全排除今年美联储降息,部分人开始押注加息。

@MiniMax_AI@minimax_aiAI 评分1717 很高兴见到 @rudrank,我们最早的模型采用者和开发者之一。更多合作即将到来😏 来 @aiDotEngineer 新加坡站和我们打个招呼吧!!🔮
引用Rudrank Riyam (@rudrank)@rudrankFinally met Leanna from @MiniMax_AI, and grateful for letting me early test the models! Eager about the next one! 👀👀
@berryxia@berryxiaAI 评分6262
引用Elon Musk (@elonmusk)@elonmuskThe latest 𝕏 algorithm has been published to GitHub github.com/xai-org/x-algorit…
@berryxia@berryxiaAI 评分2323 Gemini 3.5 Pro 的 Three.js 构建的效果。 看着挺像回事,实际效果拉不拉。 等等吧… 应该就这1/2 周



@berryxia@berryxiaAI 评分3535
引用🚨 AI News | TestingCatalog (@testingcatalog)@testingcatalogGOOGLE 🔥: New Gemini Spark screenshots featuring advanced tool use and Skills creation flow. It seems like there won't be an option to import SKILL MD files besides copeing and pasting. There is also no evidence of Browser or Computer Use atm.
@vista8@vista8AI 评分00 @vista8@vista8AI 评分1515 
@berryxia@berryxiaAI 评分4444 Qwen 3.6 Plus 和 OpenCode 免费开整啊!!!

@AYi_AInotes@ayi_ainotesAI 评分5757 

@kimmonismus@kimmonismusAI 评分2727 Codex 的“锁定使用”要来了。 大概能解释 OpenAI 昨天那张图。 “让 Codex 在你的 Mac 锁屏时也能使用”

引用🚨 AI News | TestingCatalog (@testingcatalog)@testingcatalogOpenAI is working on a dedicated setting for Codex to allow users to enable "Locked use." > Let Codex use your Mac while it's locked No more need to carry a half-open laptop around?
@MiniMax_AI@minimax_aiAI 评分2424 引用1LittleCoder💻 (@1littlecoder)@1littlecoderMinimax 🔥🔥🔥 shipping across modalities
@frxiaobei@frxiaobeiAI 评分6262 引用OpenAI (@OpenAI)@OpenAIYou've been asking for this one... Now in preview: Codex in the ChatGPT mobile app. Start new work, review outputs, steer execution, and approve next steps, all from the ChatGPT mobile app. Codex will keep running on your laptop, Mac mini, or devbox. Video
@PixVerse_@pixverse_AI 评分3131 Pixverse 让你成为全场焦点~ 在 Pixverse 网页端用 Concert Spotlight 模板制作吧! Video

@MiniMax_AI@minimax_aiAI 评分1616 在新加坡与 @zocomputer 直播!看看我们如何用 MiniMax 模型演示 Zo❤️🔥
引用Zo Computer (@zocomputer)@zocomputerKill your SaaS with Zo Computer – Live from Singapore 🇸🇬 nitter.net/i/broadcasts/1qKVmQBbk…
@alibaba_cloud@alibaba_cloudAI 评分2222 



@alibaba_cloud@alibaba_cloudAI 评分1818 
@berryxia@berryxiaAI 评分11 
@vista8@vista8AI 评分2020 正在体验朋友开发的AI工具,有点顶! 配合飞书Cli,两轮对话整理出经典AI论文合集,还配上了图表。 xiangyangqiaomu.feishu.cn/do…


@berryxia@berryxiaAI 评分2626 Violin 项目迭代后保留视频翻译多国语言的核心功能,新增用户选择目标音色、支持多角色多音色,并可在翻译成多国语言后克隆原音色,同时保留翻译后字幕导出。作者称再优化一下就能做海外视频播客了。

@vista8@vista8AI 评分3131 
@rudrank@rudrankAI 评分3030 MiniMax 模型在 ASC CLI 开发中发挥了重要作用 不知道他们有没有新模型让我试试 👀
引用Rudrank Riyam (@rudrank)@rudrankA lot of the code shipped in App Store Connect CLI has been with @MiniMax_AI M2.5 recently and the screenshot automation feature (soon) is brainstormed with it I got early access to try it out, and I am experimenting with replacing Opus 4.6 *fast*
@swyx@swyxAI 评分1313 另外我觉得现在公开披露的营收时间序列大概长这样。 预测到年底,最接近正确年底 ARR 预测的得一个赞

@AYi_AInotes@ayi_ainotesAI 评分4848 一份报告承认,中国仅用美国4%的算力就达到了后者80-90%的模型能力,报告认为这才是真正令人担忧之处。报告还指出,蒸馏攻击是真正的降维打击,几千个假账号就能低成本复制前沿模型的能力。
@AYi_AInotes@ayi_ainotes精选AI 评分6868
引用Anthropic (@AnthropicAI)@AnthropicAIWe've published a paper that explains our views on AI competition between the US and China. The US and democratic allies hold the lead in frontier AI today. Read more on what it’ll take to keep that lead: anthropic.com/research/2028-…
推荐理由:作者从商业利益角度拆解对华算力出口报告,点出NVIDIA与Anthropic在管制立场上的利益差异。
@vista8@vista8精选AI 评分6666 引用OpenBMB (@OpenBMB)@OpenBMB1/5 MiniCPM-V 4.6 (1.3B) is now live 🚀🚀 High-res visual processing, optimized for consumer-grade and mobile hardware. We’ve leveraged the latest LLaVA-UHD v4 technique to cut vision encoding costs by 55%, enabling native edge deployment with extreme efficiency. 🔥 Beats Gemma4-E2B-it and Qwen3.5-0.8B across key multimodal and Artificial Analysis benchmarks — scoring higher than Qwen3.5-0.8B using just 2.5% of its token budget. ⚡ TTFT (75.7ms) 2.2x Faster than Qwen3.5-0.8B even with 3136² high-res images. 🏗️ ~1.5x Token Throughput compared with Qwen3.5-0.8B on a single RTX 4090. Try the model here: 🤗 Hugging Face: huggingface.co/openbmb/MiniC… 💻 GitHub: github.com/OpenBMB/MiniCPM-V 🔭 Modelscope: modelscope.cn/models/OpenBMB… 🌐 Web Demo: huggingface.co/spaces/openbm… 📱 App Demo: github.com/OpenBMB/MiniCPM-V… Video
推荐理由:官方给出 1.3B 小模型处理高分辨率图像的编码成本与吞吐数据,可供端侧多模态选型参考。
@vista8@vista8AI 评分2222 有观点认为 AI 正从"宠物模式"切换到"幼儿模式":宠物靠训练、有明确边界和指令,幼儿则自己摔跤、自己爬起来、自己搞清世界如何运转。推文由此提出,世界模型加上自主改进学习的 AI 可能才是未来方向。
@vista8@vista8AI 评分3030 @berryxia@berryxiaAI 评分5454 @berryxia@berryxiaAI 评分2424 用 GPT-image-2 上传图片即可自动拆解并标注 OOTD 穿搭,提示词已放在评论区。推文以马斯克带儿子 𝕏 赴北京参会期间走红的穿搭为例演示,并附黄总吃炸酱面、志林姐姐等图片。



@berryxia@berryxiaAI 评分4545
引用Berryxia.AI (@berryxia)@berryxia这个项目也可以直接 # 安装成 Claude Code skill 命令:violin --install-skill 以后就可以直接这样:violin input.mp4 output_zh.mp4 --language Chinese 大家需要注意: 去 api.together.ai 注册获取 Key(也支持 OpenAI、ElevenLabs,只需其中一个)。 Violin 默认使用 Together AI(免费注册可得额度),需要设置环境变量: # 永久生效,加到 ~/.zshrc echo 'export TOGETHER_API_KEY=你的key' >> ~/.zshrc source ~/.zshrc
@kimmonismus@kimmonismusAI 评分00 sauce ft.com/content/9deae3c6-716d…
@kimmonismus@kimmonismus精选AI 评分8181 
推荐理由:用估值与 ARR 两组数字呈现 Anthropic 近几个月的增长速度,可作为观察头部模型公司商业化的参照。
@kimmonismus@kimmonismusAI 评分99 来源 anthropic.com/research/2028-…
@vista8@vista8AI 评分3636 vercel.com/blog/ai-gateway-p… 翻译:blog.qiaomu.ai/vercel-ai-gat…
@vista8@vista8AI 评分6464 Vercel 报告分析了 20 万个项目 7 个月共十万亿 token 的消耗数据,按费用 Anthropic 占 61% 居首,按 token 量 Google 占 38% 居首。
@kimmonismus@kimmonismusAI 评分3939 

@vista8@vista8AI 评分1616 
@MiniMax_AI@minimax_aiAI 评分3535 很高兴看到 MiniMax 在 open-multi-agent 中投入使用!🔥 它能自动将目标拆解为 DAG 任务并并行运行
引用JackChen (@JackChen_x)@JackChen_xMulti-agent's quiet problem: token cost scales with agents × turns × tool calls. It compounds fast , and that's the bill that kills production rollouts. @MiniMax_AI is now a native adapter in @JackChen_me's open-multi-agent, so goal-driven agent teams stay affordable in production, not just in demos. 👇 github.com/open-multi-agent/…