BREAKING: The results are in for Slides Arena... @AnthropicAI and @Zai_org models continue to lead the way in soft-verifiable domains 1st: Opus 4.7 by @AnthropicAI 2nd: Opus 4.7 (Thinking) by @AnthropicAI 3rd: GLM 5.1 by @Zai_org Huge congrats to @AnthropicAI and @Zai_org for establishing the SOTA for Agentic Slides
@berryxia
@berryxia · X · 历史来源 · 当前未持续收录
切换来源
@berryxia@berryxiaAI 评分3131
引用Design Arena (@Designarena)@Designarena
@berryxia@berryxiaAI 评分1313 重复造轮子的人不是傻子, 有没有一种可能只是真的是在拿AI练手和提升「熟练度」!😊 Video

@berryxia@berryxiaAI 评分3333
引用DailyPapers (@HuggingPapers)@HuggingPapersWorld Action Models: The Next Frontier in Embodied AI The first systematic survey defining WAMs as embodied foundation models that jointly predict future states and generate actions, covering architectures, data ecosystems, and evaluation protocols.
@berryxia@berryxiaAI 评分4141
引用Fred Peng (@pengzhangzhi1)@pengzhangzhi1How to Train Diffusion LLM more efficiently? Our paper has an answer for you: Don’t Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment Diffusion language models are becoming increasingly attractive: they support bidirectional generation, non-sequential decoding, and flexible editing. But training them from scratch is expensive. So a natural question is: If we already have strong pretrained autoregressive LMs, do we really need to relearn all language representations for diffusion LMs? We argue: probably not. Our view is that AR→DLM conversion should not be treated as learning language from scratch again. Much of the semantic structure is already inside the AR model. What changes is the generation order and denoising behavior. So instead of only continuing denoising training, we explicitly preserve the representation geometry of the AR model. We introduce REPR-ALIGN: during masked diffusion training, we align the hidden states of the DLM to a frozen AR teacher of the same architecture, layer by layer, using cosine similarity. No adapters. No architectural changes beyond the attention mask. Just representation alignment + masked denoising. The result: up to 4× training acceleration in our setting, with especially strong gains in low-data regimes. The main takeaway is simple: Don’t retrain the representation space from scratch. Align it, and let the model relearn the decoding path. Paper: arxiv.org/abs/2605.06885 Code: github.com/pengzhangzhi/Open… Work done with an amazing undergrad @alexisfox and advisors @Anru_Zhang @AlexanderTong7
@berryxia@berryxiaAI 评分2525 完整文章在这里: magazine.sebastianraschka.co…
@berryxia@berryxiaAI 评分5151
引用Sebastian Raschka (@rasbt)@rasbtNew article: a visual tour of recent LLM architecture advances, from Gemma 4 to DeepSeek V4. I focus on long-context efficiency tweaks like KV sharing, per-layer embeddings, layer-wise attention budgets, compressed attention, and mHC. Link: magazine.sebastianraschka.co…
@berryxia@berryxia精选AI 评分6666 Anthropic 发布内部手册《Founder's Playbook》,基于 Claude Code 和一批 YC 创始人的踩坑经验,提出 AI 会让创业失败率上升而非下降。
引用Smith铜匠・十点睡觉 (@smithandai)@smithandaix.com/i/article/205523912843…
推荐理由:Anthropic 把 Claude Code 在创业各阶段的踩坑经验拆成四个阶段,读者可对照自查原型验证与技术债。
@berryxia@berryxiaAI 评分66 
@berryxia@berryxiaAI 评分5252
引用𝐀𝐆 (@AGkorthos)@AGkorthosKorean WIRobotics just raised ~$68M. Known for its WIM wearable robots and ALLEX humanoid platform, the company plans to supply a mobile ALLEX research platform this year and target initial commercialization readiness by late 2027. Video
@berryxia@berryxiaAI 评分3535 引用GREG ISENBERG (@gregisenberg)@gregisenbergMore AI agent observations below (I keep adding to the list): 1. Hermes agents write to their own memory after every task. Which means starting today versus starting in 6 months is an unfair advantage for you. 2. We're maybe 12 months from an agent that can watch you work for a week and then do your job without any instructions. The screen recording plus agent memory plus local model combination makes this possible right now 3. The real reason local models matter for founders: you can ship a product where the AI runs entirely on the customer's device and you never touch their data. Zero privacy concerns. Zero server costs. Zero compliance headaches. That changes which industries you can sell to overnight. Healthcare, legal, finance, all the regulated verticals that won't send data to the cloud just opened up. 4. Every company needs to be rebuilt as a "second brain" before agents can be useful. That means every process, every decision, every piece of institutional knowledge has to exist in a format an agent can read. Most companies have none of this. 5. Agent costs are the new headcount. Won't be crazy for companies to spend 50%+ of their total headcount cost on tokens. 6. Agents are accidentally creating internal competition at companies. The marketing agent and the sales agent are optimizing for different metrics and working against each other without anyone realizing it. It took humans decades to develop cross-functional alignment. Nobody thought about it for agents. 7. The YAML config file is becoming the new org chart. Who reports to who, what permissions they have, what tools they access, all defined in a config file. The company's structure is literally a file you can version control, fork, and deploy. That's new. 8. The first agents that can smell a scam are going to be worth billions. Right now agents will happily wire money to a fake invoice because it matched the format. The trust layer is completely missing. 9. We're about to find out that most "expertise" was actually just memory. Knowing the tax code. Knowing the case law. Knowing which supplier charges what. When an agent holds all of that in context, the expert's value shifts from "I know things" to "I know which things matter." Much smaller group of people. 10. We're all running the same models. The differentiation is in what you feed them. Two founders with the same agent, same model, same tools will get wildly different results based purely on the quality of their knowledge base. Garbage context in, garbage output out. Forever. 11. The most underbuilt category in AI right now: agents for old people. 70 million boomers who need help with medical forms, insurance claims, and appointment scheduling. 12. Agent latency is the new page load speed. If your agent takes 45 seconds to respond, your customer already switched to one that takes 13. Skills files are the new apps. A SKILL.md that tells an agent how to do one thing well is more valuable than a SaaS subscription that does the same thing behind a login screen. 14. AI hardware... how do you create devices that are good businesses that people want? It'll be a $30 dongle you plug into existing dumb devices to give them an agent brain. Smart toaster doesn't need to be built from scratch. It needs a $30 brain attached to a $15 toaster. 15. Your agent can read faster than you can think. The bottleneck in every agent workflow is now the human approval step. We're the slow part. That's a strange thing to sit with. 16. Agents made the 80/20 rule violent. The 20% of work that matters is now the only work humans do. The 80% just disappeared. Entire job descriptions were hiding inside that 80%. 17. The thing I keep coming back to: the best businesses right now are being built by people who are just slightly ahead of their customers. Not 10 years ahead. 6 months ahead. That's the sweet spot. Far enough to lead. Close enough to be understood.
@berryxia@berryxiaAI 评分4747
引用Elliott / Shangzhe Wu (@elliottszwu)@elliottszwuCheck out Ariticraft 🦾 - a highly efficient agentic system that generates articulated 3D assets fully automatically at scale! 🚀 articraft3d.github.io/ Video
@berryxia@berryxiaAI 评分33 
@berryxia@berryxiaAI 评分5757
引用International Cyber Digest (@IntCyberDigest)@IntCyberDigest❗️🚨 BREAKING: Researchers used Mythos Preview to find the first public macOS kernel memory corruption exploit on Apple's M5 silicon, they give a glimpse into Mythos say it’s really powerful. Apple spent five years and an estimated several billion dollars building Memory Integrity Enforcement (MIE), the hardware-assisted memory safety system built around ARM's MTE. It was the flagship security feature of the M5 and A19, designed specifically to kill the entire memory corruption bug class. Researchers from Calif built a working exploit in five days. According to Apple's own research, MIE disrupts every public exploit chain against modern iOS, including the recently leaked Coruna and Darksword kits. Calif walked into Apple Park this week and handed over the report in person. Full 55-page technical report drops after Apple patches the vulnerability.
@berryxia@berryxiaAI 评分4040
引用Daily Dose of Data Science (@DailyDoseOfDS_)@DailyDoseOfDS_Transformer and Mixture of Experts, explained visually! Mixture of Experts (MoE) is a popular architecture that uses different experts to improve Transformer models. Transformer and MoE differ in the decoder block: - Transformer uses a feed-forward network. - MoE uses experts, which are feed-forward networks but smaller compared to those Transformer. During inference, a subset of experts are selected. This makes inference faster in MoE. Also, since the network has multiple decoder layers: - The text passes through different experts across layers. - The chosen experts also differ between tokens. But how does the model decide which experts should be ideal? The router does that. It is a multi-class classifier that produces softmax scores over experts to select the top K experts. The router is trained with the network, and it learns to select the best experts. But it isn't straightforward. There are challenges! Challenge 1) Notice this pattern at the start of training: - Say, the model selects "Expert 2" - This expert gets a bit better - It may get selected again since it's the "best" - It learns more - It gets selected again in the next iteration - It learns more, and so on! This means many experts can go under-trained due to the overselection of a few experts! We solve this in two steps: - Add noise to the feed-forward output of the router so that other experts can get higher logits. - Set all but the top K logits to -infinity. After softmax, these scores become zero. This way, other experts also get the opportunity to train. Challenge 2) Some experts may get exposed to more tokens than others, leading to under-trained experts. We prevent this by limiting the number of tokens an expert can process. If an expert reaches the limit, the token is passed to the next best expert. Overall, MoEs have more parameters to load. But a fraction of them are activated during inference. This leads to faster inference. Mixtral 8x7B and Llama 4 are two popular MoE-based LLMs. Have you used MoEs in production yet?
@berryxia@berryxiaAI 评分6262
引用Elon Musk (@elonmusk)@elonmuskThe latest 𝕏 algorithm has been published to GitHub github.com/xai-org/x-algorit…
@berryxia@berryxiaAI 评分2323 Gemini 3.5 Pro 的 Three.js 构建的效果。 看着挺像回事,实际效果拉不拉。 等等吧… 应该就这1/2 周



@berryxia@berryxiaAI 评分3535
引用🚨 AI News | TestingCatalog (@testingcatalog)@testingcatalogGOOGLE 🔥: New Gemini Spark screenshots featuring advanced tool use and Skills creation flow. It seems like there won't be an option to import SKILL MD files besides copeing and pasting. There is also no evidence of Browser or Computer Use atm.
@berryxia@berryxiaAI 评分4444 Qwen 3.6 Plus 和 OpenCode 免费开整啊!!!

@berryxia@berryxiaAI 评分11 
@berryxia@berryxiaAI 评分2626 Violin 项目迭代后保留视频翻译多国语言的核心功能,新增用户选择目标音色、支持多角色多音色,并可在翻译成多国语言后克隆原音色,同时保留翻译后字幕导出。作者称再优化一下就能做海外视频播客了。

@berryxia@berryxiaAI 评分5454 @berryxia@berryxiaAI 评分2424 用 GPT-image-2 上传图片即可自动拆解并标注 OOTD 穿搭,提示词已放在评论区。推文以马斯克带儿子 𝕏 赴北京参会期间走红的穿搭为例演示,并附黄总吃炸酱面、志林姐姐等图片。



@berryxia@berryxiaAI 评分4545
引用Berryxia.AI (@berryxia)@berryxia这个项目也可以直接 # 安装成 Claude Code skill 命令:violin --install-skill 以后就可以直接这样:violin input.mp4 output_zh.mp4 --language Chinese 大家需要注意: 去 api.together.ai 注册获取 Key(也支持 OpenAI、ElevenLabs,只需其中一个)。 Violin 默认使用 Together AI(免费注册可得额度),需要设置环境变量: # 永久生效,加到 ~/.zshrc echo 'export TOGETHER_API_KEY=你的key' >> ~/.zshrc source ~/.zshrc
@berryxia@berryxiaAI 评分44 项目完整介绍在这里:media.mit.edu/projects/elect…
@berryxia@berryxiaAI 评分4242
引用Space and Technology (@spaceandtech_)@spaceandtech_MIT researchers have developed new artificial muscles called Electrofluidic Fiber Muscles for robots and wearable devices. These flexible muscles can be woven into fabric and work silently without bulky equipment. The system is lightweight, portable, and uses tiny fiber pumps smaller than 2 millimeters to generate powerful movement directly from electricity. Video
@berryxia@berryxiaAI 评分4242 
引用Berryxia.AI (@berryxia)@berryxia关于Claude 封号,如何申请美区退款! 这件事,我给大家简单交代一下后续。 因为我当时订阅是用 Apple Gift Card 礼品卡充值的,所以它没有自动退费。 我订阅的是 Max 125 美金那一档。 我刚刚给苹果中国打了电话,具体操作流程如下: 1. 拨打 Apple Store 对应的 400 电话,客服会进行初步了解。 2. 提供你的 Apple ID。 3. 随后电话会转接到外区同事。虽然是外区,但讲普通话也没问题(我接通的是台湾同事,中文沟通很顺畅)。 客服会提供两种退款方式: 1. 到网页上自主申请退款。 2. 直接告诉客服,让他帮你手动退款。你只需要确认 Apple ID 和对应金额,他就会帮你提交申请。 退款一般会在 48 小时内原路退回。 如果大家有被封号且没有收到自动退款的,可以尝试这样操作。
@berryxia@berryxiaAI 评分6363 引用Berryxia.AI (@berryxia)@berryxia兄弟们,这个可以啊!赶紧装起来! Kevin Lin,牛津大学博士后,前Meta和Microsoft研究员,刚刚把Violin这个开源视频翻译Skill放了出来。 视频已经是互联网绝对主流的内容形式。 可绝大多数高质量讲座、演讲、播客却被单一语言死死锁住,全球观众根本触达不到。 Violin把ASR、LLM翻译、TTS三者无缝串成一条流水线。 「输入一段视频,它就能自动完成语音识别、多语言翻译、自然语音合成。」 最实用的是两个功能: 你可以个性化翻译风格,把学术报告改成孩子也能听懂的版本; 还能直接和视频聊天,任何问题都基于视频内容给出答案。 它同时支持Web应用、CLI命令行和Agent Skill,全部MIT开源。 以后高质量内容不再只属于某一种语言,而是真正走向全球。 Demo、博客和GitHub都在原帖。 如果你在做内容、教育、跨语言传播,或者正在开发多模态Agent,这套Skill值得立刻去试。 你觉得AI下一步最该解决的,是内容创作,还是内容全球化? 项目地址:github.com/shang-zhu/violin Video
@berryxia@berryxiaAI 评分5454
引用Kevin Lin (@KevinQHLin)@KevinQHLin🌟Introducing🎻Violin — an Open-source Video Translation Skill. 📹Video is the dominant medium on the internet, yet most high-quality content (lecture, talk, podcast) is locked behind a single language, leaving global audiences behind. So we built Violin: a video skill that combines speech recognition, LLM translation, and speech synthesis into one seamless pipeline. 🌐 Demo: violin-ai.com 📝 Blog: together.ai/blog/violin-open… 🔗 GitHub: github.com/shang-zhu/violin ✨Key Features: 🎙️High-quality multilingual ASR & Translation & TTS. 🗣️Personalize translation & voice (turn an academic talk into something children can follow). 💬Chat with the video — ask any questions grounded in the video. 🧩Support Web app, CLI, and Agent skill 🍃Fully open-source under MIT. ❤️Built with the wonderful @ShangZhu18 and advised by @james_y_zou ! All features powered by @togethercompute . Try it and let us know what you think! 🎻 Video
@berryxia@berryxiaAI 评分44 对对对! 这个就是和JigSpace 的发动机的设计很接近的效果,磨具在细化一下就更好了。 Video

@berryxia@berryxiaAI 评分5858 开发者 neilsonks 开源了一个面向 Claude Code 的 3D 生成工具包 image-blaster,输入一张图片即可生成包含环境、网格、物理、灯光和音频的可交互 3D 世界。
引用neilson (@neilsonks)@neilsonksopen-sourcing a 3D gen toolkit for Claude Code input image → environment, meshes, physics, lighting, & audio Video
@berryxia@berryxiaAI 评分2121 @berryxia@berryxia精选AI 评分6666
引用Prime Intellect (@PrimeIntellect)@PrimeIntellectAutomating AI research is the next major step in AI We let Claude Code (Opus 4.7) and Codex (GPT 5.5) run autonomously on the nanoGPT speedrun optimizer track using our idle compute. ~10k runs, ~14k H200 hours Opus now holds the record at 2930 steps vs the 2990 human baseline
推荐理由:Prime Intellect 用闲置算力让智能体自主优化 nanoGPT 训练,显示其擅长组合已有方法但在创新上受限,实验日志已开源。
@berryxia@berryxiaAI 评分6161 引用yetone (@yetone)@yetone由于这篇文章太伟大了,所以我把它变成了一个 Agent Skill。 大家可以使用自己的 Coding Agent 安装一下这个 Skill,这样就可以用「最佳实践」来轻松地重构或者开发一个既容易跨平台、又极其接近 Native 性能的桌面端应用。 github.com/yetone/native-fee…
@berryxia@berryxiaAI 评分5353 
@berryxia@berryxiaAI 评分6363
引用Anthropic (@AnthropicAI)@AnthropicAIWe've published a paper that explains our views on AI competition between the US and China. The US and democratic allies hold the lead in frontier AI today. Read more on what it’ll take to keep that lead: anthropic.com/research/2028-…
@berryxia@berryxiaAI 评分3535 LM Studio 又更新了 Beta 版,在 MLX 框架下优化增强了之前的缓存问题。 目前需要打开 dev 模式然后加油更新到最新版体验。 Video

@berryxia@berryxiaAI 评分5858 Codex 移动手机版已经上线,可直接在应用商店下载使用,iOS 端已经用上。安卓端可以到 Google Play 查看是否有该应用。

@berryxia@berryxiaAI 评分6060 引用Roberto Nickson (@rpnickson)@rpnicksonMeta just launched Incognito Chat with Meta AI - the world's first truly private way to chat with AI. But I had a lot of questions. How does it actually work? Isn't this contradictory to Meta's business model? What assurances do we have that this is airtight? I had a chance to speak to the Head of Whatsapp @wcathcart and Meta's VP of AI Products @vishalshahis to learn more. Video
@berryxia@berryxiaAI 评分77 @berryxia@berryxiaAI 评分6161 引用Tom Huang (@tuturetom)@tuturetom正式开源 html-anything 🚀 1:1 让你感受全网爆火 Claude code 作者提的 HTML 效果! 你的 Agent 现在可以将任何数据转为世界级设计水准的 HTML 🔥 历时 3 天,1万五千行代码!支持 75 套 Skills,9 种导出格式,支持所有的 code agent,包括 claude code、codex、openclaw、hermes 等💥地址见评论区 Video