@berryxia
@berryxia · X · 历史来源 · 当前未持续收录
切换来源
@berryxia@berryxiaAI 评分88 
@berryxia@berryxia精选AI 评分7979
引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlysGoogle’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @GoogleDeepMind gave us pre-release access to Gemini 3.5 Flash, the latest model in its Flash family, which has traditionally has offered faster, lower-cost alternatives to Gemini Pro models. Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash, driven primarily by agentic performance gains and hallucination reduction. It achieves speeds of over 280 output tokens/s, but higher token usage and token pricing make it over 5x more costly to run the Intelligence Index than Gemini 3 Flash, and 75% more costly than Gemini 3.1 Pro. Gemini 3.5 Flash is $1.50/1M input and $9/1M output tokens, Gemini 3 Flash was $0.5/$3 per 1M input/output tokens, a 3x increase. The rest of the increase was driven by higher token usage when running our benchmarks Key results for Gemini 3.5 Flash with ‘high’ thinking level: ➤ 9 point Intelligence Index improvement: Gemini 3.5 Flash scores 55 on the Artificial Analysis Intelligence Index, up 9 points from Gemini 3 Flash. This places it ahead of Grok 4.3 (high, 53) and Claude Sonnet 4.6 (max, 52). The model improves across nearly all evaluations, with the largest gains coming from agentic evaluations and AA-Omniscience (knowledge and hallucination). On AA-Omniscience, Gemini 3.5 Flash improves by 11 points, driven primarily by reduced hallucinations, with its hallucination rate falling to 61%, a 31 point decrease compared to Gemini 3 Flash ➤ Agentic capability improvements: Gemini 3.5 Flash improves substantially over Gemini 3 Flash across our agentic evaluations, in both GDPval-AA (real-world agentic tasks) and Tau2-Bench Telecom (agentic tool use). Its GDPval-AA result is especially notable, achieving an Elo of 1656, well ahead of Gemini 3 Flash (1204) and Gemini 3.1 Pro (1314), and just behind GPT-5.4 (xhigh, 1674). This represents a meaningful step forward for Google in agentic performance, which has historically been a relative weakness for Gemini models ➤ Speed-intelligence frontier: Gemini 3.5 Flash achieves speeds of over 280 output tokens per second, ~70% faster than Gemini 3 Flash and models such as gpt-oss-120b and GPT-5.4 mini (xhigh). With its 55 Intelligence Index score, this places Gemini 3.5 Flash on the speed-intelligence Pareto frontier alongside Gemini 3.1 Pro and Gemini 3.1 Flash-Lite, reinforcing Google’s strength in models balancing speed and intelligence ➤ 5.5x increase in cost to run: Gemini 3.5 Flash costs $1,552 to run the Artificial Analysis Intelligence Index, 5.5x more than Gemini 3 Flash and 75% more than Gemini 3.1 Pro. This is driven by increases in both token usage and token prices. Output token usage is broadly unchanged from Gemini 3 Flash (73M vs. 72M), but input token usage increases significantly, driven primarily by an increase in the number of turns in agentic evaluations. Gemini 3.5 Flash is priced 3x higher than Gemini 3 Flash at $1.50/$9.00 per 1M input/output tokens, with a 90% discount for cached input tokens ➤ Google continues to lead multimodal performance: Gemini 3.5 Flash is multimodal, supporting image, video, and speech input alongside text. This differs from many proprietary models, including Claude Opus 4.7, Grok 4.3, and GPT-5.5, which support image input only. In our multimodal evaluation, MMMU-Pro, Gemini 3.5 Flash scores 84% - the highest score recorded. This puts models from Google in the top two spots, with Gemini 3.1 Pro scoring 82% Key model details: ➤ Context window: Retains the same 1M context window as Gemini 3 Flash ➤ Multimodality: Text, image, video and speech input with text output only ➤ Pricing: $1.50/$9.00 per million input/output tokens, with a 90% discount for cached input tokens Congratulations @GoogleDeepMind , @sundarpichai and @demishassabis on the great release!
推荐理由:借 Artificial Analysis 的预发布基准,可以看到 Gemini 3.5 Flash 在智能与速度上的提升及其成本代价。
@berryxia@berryxiaAI 评分1616 作者提醒做AI不要忽视身体健康,AI迭代永远追不完。普通人除了豆包这类chatbot,还可以用AI工具打造专属健身教练,辅助自己更好地锻炼。
引用Berryxia.AI (@berryxia)@berryxiax.com/i/article/205664131387…
@berryxia@berryxiaAI 评分88 我是想给一些零售连锁品牌做一些类似的实时模型的分析,或者也不用实时,进行视频解析就够了。 未来识别会看到更多的类似的场景应用啊~ Video

@berryxia@berryxiaAI 评分22 不是,兄弟们。😂 是我孤陋寡闻了么? 还是瓦特了…… 我一个非B站用户,好久没有打开查个东西。 现在B站的风格已经如此奔放了……

@berryxia@berryxia精选AI 评分7676 
推荐理由:归纳了 Gemini 3.5 系列、Omni 世界模型与硬件落地的要点,可据此了解 Google 在智能体方向的推进节奏。
@berryxia@berryxia精选AI 评分7575
引用Google DeepMind (@GoogleDeepMind)@GoogleDeepMindWe’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video. It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing 🧵 Video
推荐理由:Gemini Omni 把生成视频做成可对话编辑的对象,并同步在 Gemini App 等入口上线,读者可据此观察视频生成向可编辑素材演进。
@berryxia@berryxiaAI 评分2727 Gemini 3.5 flash 使用反重力工具,一句话使用多个 Agent 同时写作构建整个城市的过程,还挺有意思的。 Video

@berryxia@berryxia精选AI 评分7373 
推荐理由:材料交代了 Gemini Omni 面向订阅层的开放节奏与视频优先的输出形态,读者可据此判断上手门槛。
@berryxia@berryxiaAI 评分2828 
@berryxia@berryxiaAI 评分3434 还不知道在哪里看直播的兄弟们注意了! Google I/O 直播观看地址👇 #google
引用Google Gemini (@GeminiApp)@GeminiAppIt’s #GoogleIO Day One. Who’s ready to see what’s coming to Gemini? Livestream starts here at 10am PT: nitter.net/i/events/2053241348807…
@berryxia@berryxiaAI 评分5151
引用Elon Musk (@elonmusk)@elonmuskAnthropic will not be destroyed. Their AI+harness goes far beyond coding and Opus 4.7 is still better than Composer 2.5, albeit a lot more expensive. Cursor is however an important piece of the puzzle to make Grok much better.
@berryxia@berryxiaAI 评分1818 引用烟花老师 (@teach_fireworks)@teach_fireworks还有一百多就五千订阅了,不知道一觉醒来会不会有惊喜。我经常不按常理出牌,就提前写好庆祝5k订阅达成吧,哈哈🎆 我主业是一个AI架构师,也是一支烟花AI社区的联创,从23年至今大概积累了40个垂直的AI社群,大家都很纯粹 全都是免费的社群,基本上都是研发,产品和创业者和行业大佬,也欢迎大家一起进群交流,可以在这里登记信息,我会邀请大家进群 (长期有效) hqexj12b0g.feishu.cn/share/b… 虽然X算法偶尔抽风 还会误杀,不可否认X上的算法还算公平的,之前我分别在公众号,小红书和抖音尝试了蛮久自媒体,也是差不多的输出,最终还是X上正反馈更多一些,其他的平台都一言难尽。 争取今年做到一万粉。谢谢订阅我的朋友们,以后继续输出更多干货! 不过相当于X的收获,最大的惊喜是之前开源的fireworks-tech-graph 快7k star 了,靠神佬等众多大佬的喜爱转发,基本全靠X平台的传播,也合并了不少PR,我基本上没有在国内自媒体宣传过这个项目,完全靠自来水推荐,非常幸运可以感受到了流量加持后项目开源。 不过流量来得快去的也快,我内心也算比较平静,大大小小写了快20个开源项目,由于我懒得宣传,基本上是我自己在用,有几个harness 相关的skill 真的不错,大家可以去看下,总有一款你喜欢。
@berryxia@berryxiaAI 评分6363 @berryxia@berryxia精选AI 评分6565 引用Yukang Chen (@yukangchen_)@yukangchen_🚀 Excited to release LongLive 2.0! 🎬 An end-to-end infrastructure for long video generation, with FP4 and parallelism at the core of both training and inference. ⚡45.7 FPS generation speed on 5B model⚡ ✨ LongLive 2.0 supports real-video training, few-step distillation, multi-shot training/inference, sequence-parallel acceleration, NVFP4 KV cache, and async VAE decoding deployment. 🧩 To our knowledge, this is the first open-source 4-bit long video generation infra that covers both training and inference. 🙌 Welcome to check it out, try it, and share feedback! 🔗 Code: github.com/NVlabs/LongLive 📰 Paper: huggingface.co/papers/2605.1… 🎥 Demo: nvlabs.github.io/LongLive/Lo… #LongVideoGeneration #VideoGeneration #Realtime #AIInfra #EfficientAI #FP4 #Parallel #NVIDIA Video
推荐理由:开源方案把 FP4 量化与并行加速同时用在训练和推理,读者可据此了解长视频实时生成的技术路线。
@berryxia@berryxiaAI 评分2020
@berryxia@berryxiaAI 评分2424 引用Business (@XBusiness)@XBusinessx.com/i/article/205536179511…
@berryxia@berryxiaAI 评分2626 斯坦福数学家 George Pólya 用40年观察发现,聪明学生卡在难题上并非因为笨,而是没人教他们动手前该做什么——问题一出现就焦虑地立刻开算,越努力越偏。
引用Dr.Xiao.AI (@xiaoxiao_2580)@xiaoxiao_2580A Stanford mathematician spent forty years watching one brilliant student after another crash into hard problems. Not because they weren’t smart. But because no one had ever taught them what to actually do “before” they started solving. His name was George Pólya. In 1945, he published “How to Solve It”. The book sold over a million copies and has never gone out of print. Even Marvin Minsky, who built the first neural network, said that everyone should read it. Yet most people still haven’t heard of it. What Pólya kept seeing was the same failure pattern, again and again: The moment a difficult problem appears, students get anxious and immediately start calculating. Not because calculating is the right first step, but because doing somethingfeels much more comfortable than sitting with “I don’t know.” They end up working hard in completely the wrong direction. The step he found most neglected was this: Truly understand the problem first. Not just skim it. Not just think “this looks familiar.” His test was simple but ruthless: Can you restate the problem in your own words without looking at the original? If you can’t, you don’t actually understand it yet. Most people skip this step entirely. They jump straight into execution and then get stuck on a problem they never truly grasped. Pólya outlined four steps for solving problems. But in real life, the two that matter most are usually the first and the last: 1. Understand the problem deeply 2. Devise a plan (if you’re stuck, try solving a simpler version first and bring the insight back) 3. Carry out the plan 4. Look back — verify, generalize, and reflect The people who get truly good at this aren’t the ones who practice more. They’re the ones who’ve learned to slow down when every instinct is screaming at them to just start calculating — especially at the beginning, and again at the end. What struck me most after reading this: We assume hard problems are difficult because they are. Most of the time, it’s simply because we never took the time to truly understand them.
@berryxia@berryxia精选AI 评分8080 引用Andrej Karpathy (@karpathy)@karpathyPersonal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
推荐理由:Karpathy 宣布加入 Anthropic 并重返一线 R&D,他此前是 OpenAI 创始团队成员和 Tesla AI 前总监。
@berryxia@berryxia精选AI 评分7878 引用Andrej Karpathy (@karpathy)@karpathyPersonal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
推荐理由:Karpathy 从教育与顾问角色重返前沿实验室,为观察大模型竞争的人才流向提供了一个新样本。
@berryxia@berryxiaAI 评分3535 马斯克(Elon Musk)直接回复"True",认同黄仁勋关于 AI 的观点。黄仁勋称 AI 只会取代重复机械劳动,让人专注于更有意义的工作,同时把全球 GDP 从 100 万亿推到几百万亿。

@berryxia@berryxiaAI 评分2626 引用Lucius (@LuciusHQ)@LuciusHQWe raised $3M to build Lucius AI - the Context Layer for Your Organization. Backed by Future Capital Discovery Fund, we’re tackling a problem we kept running into ourselves: Individuals ship 10× faster with AI. Organizations don't. Over 30% of your team's time is spent rebuilding context someone already had. It shows up everywhere a decision was already made but can't be found again - community operations, customer support, pre-sales reception, sales research, project management, internal collaboration. We're building Lucius to close that gap. Video
@berryxia@berryxiaAI 评分3232 引用向阳乔木 (@vista8)@vista8小龙虾和Hermes的热度在AI科技圈终于降了。 按扩散发展规律看,民间热度估计刚刚开始。 但对于普通用户来说,各种龙虾类Agent产品上手难度还是有点高。 如提示词怎么写、工作流怎么配、模型怎么选,全靠自己摸索。 不少大厂提供了 OpenClaw 和Hermes的安装镜像。 但折腾起来依然费劲,普通人才懒得研究。 结果就是:会用的人锦上添花,不会用的人装完吃灰。 普通用户最需要的是开箱即用的龙虾产品。 不得不说,360还是很懂普通用户痛点,开发了360安全龙虾云端版。 内置了100+预训练好的专家虾,不用自己从零调教,对应场景拿来就用,接入大量语言模型、生图模型和视频生成模型,搭配技能市场,能把各种常见工作流都串起来。 而且手机上也能用,随时对话调教优化。 甚至还准备了个「龙虾教练」,解决普通"不会训龙虾"的问题,让它带着走,10分钟能训出一只针对自己场景的专属虾。 下载安装地址 claw.360.cn,全平台都有,微信、飞书、钉钉也能接入,感兴趣可以试试。
@berryxia@berryxiaAI 评分3131 我就想知道小Happy这个月的工资可以拿到手吗? 岂不是token爆炸了哈哈 ~~ 同情小编几秒钟~~
引用Happycapy (@happycapyai)@happycapyaiI can control my Mac with hapoycapy! Connect Your Mac in 3 Steps Step 1: Open Terminal on your Mac Press `Cmd + Space`, type `Terminal`, hit Enter. Step 2: Paste and run this command curl -fsSL 'happycapy.ai/api/bridge/inst…' Step 3: Done! Your Mac is now connected. Tell me what you want to do. Video
@berryxia@berryxiaAI 评分2626 . @dangreenheck 老哥这个原版看着还是最牛逼! 我的还是太潦草哈哈,我得值40美金不😄 Video
引用Berryxia.AI (@berryxia)@berryxia我靠!我又行了啊,兄弟们~ 真的是Saas 已死,Agent 称王的时代来了 !! 我今天花了2小时,就用Cursor + Claude把海外老哥卖149美元的「Three.js热带海洋实时交互系统」直接手搓复刻出来了。 😄 实时交互全都有:海洋波浪动态、风速实时调节、天空环境光变化…… 一整套物理交互。 原版我不知道实际交付效果如何,但我这个版本视觉和交互已经还原80%以上,还额外加了中英文双语切换、海洋动植物实时互动、更多细节物理反馈。 这个思路还能往天气系统、生态模拟、甚至教育场景里疯狂扩展。 以前要花149美元买的东西,现在AI两小时就能自己造出来。 感兴趣的朋友点赞破100,我就直接把完整代码开源给大家玩! 破不了就算了…… 我消耗的token已经够我心疼的了哈哈。 (附上我现在跑起来的实时演示效果图/视频) 原系统项目见评论区下👇🏻: Video
@berryxia@berryxiaAI 评分3434
引用Alberto (@SwiftyAlbert)@SwiftyAlbertPor favor mirad qué maravilla el trazado cefalométrico asistido por IA, aunque no entendáis de ortodoncia: Video
@berryxia@berryxiaAI 评分22 
@berryxia@berryxiaAI 评分2929
引用Berryxia.AI (@berryxia)@berryxia我靠!不是,我是最后一个知道的吗??? 你们的嘴可真严啊,Cursor选择Auto模式下。 居然不需要魔法网络就可以使用啊!
@berryxia@berryxiaAI 评分77 @berryxia@berryxiaAI 评分4242
引用Tencent Hy (@TencentHunyuan)@TencentHunyuan🎉 🎉 🎉 We're open-sourcing Chronicles-OCR, a visual perception benchmark evaluating VLLMs on ancient Chinese characters. The dataset spans 3,000 years of evolution. It covers 7 historical scripts from Oracle Bone to Cursive, featuring 2,800 balanced images across highly diverse physical media. We assess models on 4 core tasks: • Character Spotting • Fine-grained Recognition • Ancient Text Parsing • Script Classification The evaluation reveals how visual distribution shifts affect model perception over time. Explore the dataset and paper below. 👇 📄 Paper: arxiv.org/abs/2605.11960 🔗 GitHub: github.com/VirtualLUOUCAS/Ch…
@berryxia@berryxiaAI 评分3030 我靠!不是,我是最后一个知道的吗??? 你们的嘴可真严啊,Cursor选择Auto模式下。 居然不需要魔法网络就可以使用啊!

@berryxia@berryxiaAI 评分3737 今晚有 Google I/O 大会,是时候拿出来点新东西了啊。 不然 A 社、O 社最近都有点狂妄了啊~ Google 老哥加油啊~💪🏻
引用Google Gemini (@GeminiApp)@GeminiAppOur biggest event of the year is almost here and we've got a virtual front-row seat saved just for you. Join us to watch the 2026 #GoogleIO livestream tomorrow (5/19) at 10am PT: nitter.net/i/events/2053241348807…
@berryxia@berryxiaAI 评分2121 Gemini 3.2 测试视频,这个看着还是有点东西啊! 实际到底如何,我们等 Google 大善人发布。 Video

@berryxia@berryxiaAI 评分2020 有人把 Apple Vision Pro 的眼球追踪交互方式搬到了普通显示器上:眼球移动充当鼠标指针,手部捏合动作充当鼠标点击。推文认为这套方案用在显示器上"有点意思",并附有一段视频演示。

@berryxia@berryxiaAI 评分4040
引用Vivek Sen (@Vivek4real_)@Vivek4real_JENSEN HUANG: “IF I HAVE A CHOICE BETWEEN A NEW COLLEGE GRADUATE WITH NO CLUE WHAT AI IS AND ONE THAT IS EXPERT IN USING AI, I WOULD HIRE THE ONE WHO'S EXPERT IN USING AI. ACCOUNTANT, MARKETING, SUPPLY CHAIN, LAWYER, SALESPERSON. EVERY SINGLE TIME.” Video
@berryxia@berryxiaAI 评分1919 Gemini 视频 Veo4.𝕏 ? 要来了,期待超越 SD2 啊!兄弟们~~
引用Logan Kilpatrick (@OfficialLoganK)@OfficialLoganKGemini Video
@berryxia@berryxiaAI 评分5757
引用Cursor (@cursor_ai)@cursor_aiIntroducing Composer 2.5, our most powerful model yet. It's more intelligent, better at sustained work on long-running tasks, and more reliable at following complex instructions. For the next week, we’re doubling the included usage of the model.
@berryxia@berryxia精选AI 评分6666
引用Odyssey (@odysseyml)@odysseymlIntroducing Agora-1, a multi-agent world model. Multiple participants—human or AI—can now interact inside the same world simulation, all in real-time. Try our playable research preview today, with Agora-1 simulating a multiplayer GoldenEye deathmatch! Video
推荐理由:世界模型从单人视频生成扩展到多人实时共享模拟,读者可据此了解人机共处同一模拟世界的当前形态。
@berryxia@berryxiaAI 评分3737
引用Odyssey (@odysseyml)@odysseymlMeet our new friend, Starchild-1 ❤️ Starchild-1 is the first ever real-time multimodal world model. A world model understands and simulates the world. Starchild-1 has learned to generate not just the visuals of the world, but the sounds of it too! Video
@berryxia@berryxia精选AI 评分6868 引用Anthropic (@AnthropicAI)@AnthropicAIAnthropic is acquiring @stainlessapi, an SDK and MCP server platform that has powered every Anthropic SDK since the earliest days of our API. Read more: anthropic.com/news/anthropic…
推荐理由:作者把 SDK 与 MCP 工具链的归属变化作为切口,讨论这次收购会如何影响开发者的选择空间。