X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(541)
@fofrAI@fofrAIAI 评分99 @ArtificialAnlys@ArtificialAnlysAI 评分2121 
@ArtificialAnlys@ArtificialAnlysAI 评分1515 
@ArtificialAnlys@ArtificialAnlysAI 评分4949 
@kimmonismus@kimmonismusAI 评分22 @kimmonismus@kimmonismusAI 评分4444
引用@kimmonismus@kimmonismusWTF I was on a plane for 10 hours. Within 10 hours OpenAI released a blogpost that their newly trained model solved 100(!!) open long standing math problems that even surprised their own resecaher. MiMo released an amazing model and Grok 4.7 underperformed. I’ll never take a plane again. Ever. (Ok, not true).
@cb_doge@cb_dogeAI 评分55 抱歉,主推文内容仅包含一个链接(https://t.co/Fw1esTkegn),没有可翻译的文字内容。请提供推文正文文本,我将为您翻译。

@fchollet@fcholletAI 评分2626 @fchollet@fcholletAI 评分3232 @SemiAnalysis_@SemiAnalysis_AI 评分3030 
@SemiAnalysis_@SemiAnalysis_AI 评分1919 这些就是我们在 ChipBook 中追踪的数据集类型,用来在公告发布之前看清货物流向。点击这里查看。👇️(3/3) https://t.co/geycIFHBmg
@kimmonismus@kimmonismusAI 评分2727
引用@kimmonismus@kimmonismusI was too quick and excited when landed in SF. As most of you pointed out the real deal is actually MiMo-2.6! Better than Grok 4.7, much cheaper and holy moly is china back. That is the real surprise! I need to check that out. Message to chubby: next time checking x properly before posting :D
@fchollet@fcholletAI 评分99 @kimmonismus@kimmonismusAI 评分1515 @kimmonismus@kimmonismusAI 评分66 @swyx@swyxAI 评分1010 @latentspacepod @allenpark 现已上线,凡是有优质播客的地方都能听到! (邮件、Apple Podcasts、YouTube) https://t.co/AMwZCJ5IEK
@EMostaque@EMostaqueAI 评分99 @kimmonismus@kimmonismusAI 评分2424 是的,你们说得对:https://t.co/NFAgw6y2rv
引用@kimmonismus@kimmonismusI was too quick and excited when landed in SF. As most of you pointed out the real deal is actually MiMo-2.6! Better than Grok 4.7, much cheaper and holy moly is china back. That is the real surprise! I need to check that out. Message to chubby: next time checking x properly before posting :D
@kimmonismus@kimmonismusAI 评分2525 

引用@kimmonismus@kimmonismusDamn. I’m late to the party. Plane just arrived in SF: Grok 4.7 looks super amazing! Not only is it competitive to sota and even surpasses fable 5.1 in some benchmarks, it’s also so much cheaper than GPT and Fable 5.1! Well done spaceX!! Will write about it a bit more later on https://t.co/MADklQFtIs https://t.co/gYtzugfONx
@testingcatalog@testingcatalogAI 评分4646 
引用@finkd@finkdTeaming up with Shopify to make shopping and checkout easier in Muse. Shoppers find more. Shops sell more. More partnerships like this coming soon. https://t.co/ccak4J7IIb
@ClementDelangue@ClementDelangueAI 评分1414 训练模型正变得越来越容易——看看这个和 TRL——尤其是配合智能体! 如果你所有任务还在用现成的模型,那你可就落伍了!https://t.co/ZAoGLIPb6A
@kimmonismus@kimmonismusAI 评分22 @EMostaque@EMostaqueAI 评分4343 没有哪家 AI 模型公司会理性地放慢脚步,除非他们拥有一个能解决 100 个长期未解数学难题的模型 https://t.co/nazevM7O0O https://t.co/b7qVHAP0fn
引用@OpenAI@OpenAIWe’re working with an independent advisory group of mathematicians to help OpenAI responsibly share advances in AI and mathematics. The group will advise on how we assess and communicate new mathematical results, uphold academic and professional standards, and build tools that support mathematical research and learning. Through this work, we want mathematicians to be at the center of shaping how AI supports mathematical understanding and how its benefits reach the wider community. https://t.co/QCMFLFlIv8
@ericzakariasson@ericzakariassonAI 评分44 @ericzakariasson@ericzakariassonAI 评分66 @ericzakariasson@ericzakariassonAI 评分3131 knowledge work https://t.co/yuxNpjBFJV
引用@cb_doge@cb_dogeBREAKING: Grok 4.7 beats GPT-6 Astra in both real-world work benchmarks: • Professional knowledge work: 1,695 vs 1,542 • Multi-hour office work: 1,657 vs 1,569 Grok 4.7 is the clear winner across both. 🔥 https://t.co/2do2x53THs
@ericzakariasson@ericzakariassonAI 评分66
eric zakariasson@ericzakariassonAI 评分3232引用aditya@adxtyahqGROK 4.7 IS ACTUALLY COMPETING WITH GPT-6 ASTRA. I gave GPT-6 Astra, Grok 4.7, Kimi K3 and Fable 5.1 the same prompt to build a flight simulator Astra was still #1 overall, but Grok 4.7 was surprisingly close Kimi K3 and Fable 5.1 were basically a draw and both produced a much smoother result, while Grok 4.7 was right up there with Astra in terms of overall quality. overall: Astra > Grok ≈ Kimi ≈ Fable GPT-6 Astra finally has some serious competition.
Logan Kilpatrick@OfficialLoganKAI 评分3434如果你在用 AI 构建产品,你应该花超过 25% 的时间做基准测试,并努力让模型实验室关注这些基准测试 这是加速公司进展的最简单路径
@testingcatalog@testingcatalogAI 评分5454 

@kimmonismus@kimmonismusAI 评分4646
引用@SpaceXAI@SpaceXAIGrok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed. https://t.co/H3OTBbXyvO
@natolambert@natolambertAI 评分2020 强烈赞同。这种情况比我预期的要少得多——部分原因是构建高质量环境通常涉及相当昂贵的验证(测试强模型)。不过,还可以有更多。
引用@Thom_Wolf@Thom_Wolfreleasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now the equivalent of sharing high quality pretraining data but in the new RLVR paradigm https://t.co/Sg2mYfwswI
@testingcatalog@testingcatalogAI 评分1616 @testingcatalog@testingcatalogAI 评分3030 
@dongxi_nlp@dongxi_nlpAI 评分3333 小米 MiMo-V2.6 30 steps,训练成本 350万美元! https://t.co/Akdfm8dQSs https://t.co/tbZBD9Ny7C

@openclaw@openclawAI 评分33 https://t.co/WrXxjd4Qf3 https://t.co/aOEcvGFcf8

@emollick@emollickAI 评分2222 (我的意思是,AI 公司不会被允许简单地替代医生和律师,而是需要协商它们在一个充满重叠的社会与职业关系和义务的世界中如何运作——这些关系和义务超越了任务本身,正如它们在数学领域发现必须做的那样)
@opencode@opencodeAI 评分3434 @OpenRouter@OpenRouterAI 评分4242 @OpenRouter@OpenRouterAI 评分5959