跳到正文

X:Elvis Saravia

@omarsar0 · X

当前显示全部 AI 相关新闻
切换来源
全部X新闻X:Rohan Paul1130 条X:Kim1108 条X:阿易 AI Notes821 条X:阿里云 / Alibaba Cloud415 条X:Testing Catalog328 条X:Elvis Saravia312 条X:OpenRouter287 条X:Elon Musk285 条X:cb_doge276 条X:Artificial Analysis257 条X:Ethan Mollick241 条X:PixVerse235 条X:Replit213 条X:OpenAI Developers199 条X:SemiAnalysis190 条X:ZHO182 条X:Alexandr Wang(Scale AI 创始人/Meta 首席 AI 官)173 条X:小北161 条X:swyx161 条X:Gemini156 条X:MiniMax154 条X:Tibo122 条X:OpenAI119 条X:Dex Horthy(HumanLayer)112 条X:Google AI for Developers103 条X:X.PIN100 条X:蚂蚁百灵98 条X:Emad Mostaque96 条X:Thomas Wolf(Hugging Face 联创/CSO)93 条X:Epoch AI91 条X:Luma AI88 条X:AI Safety Memes87 条X:Claude Devs84 条X:Runway81 条X:Aravind Srinivas(Perplexity CEO)80 条X:Jason Liu80 条X:马东锡 NLP77 条X:面壁智能 OpenBMB77 条X:Sam Altman77 条X:Frank Wang 玉伯73 条X:fofr70 条X:Gabriel70 条X:洪明68 条X:阑夕61 条X:opencode59 条X:Eric Zakariasson58 条X:Francois Chollet58 条X:Perplexity58 条X:Yuchen Jin58 条X:Greg Brockman57 条X:赵纯想55 条X:Nathan Lambert55 条X:Tencent WorkBuddy55 条X:Microsoft Research54 条X:Cohere51 条X:腾讯混元50 条X:Suno49 条X:Claude45 条X:Anthropic44 条X:Clément Delangue(Hugging Face CEO)44 条X:Peter Steinberger44 条X:Krea AI41 条X:OpenClaw41 条X:通义千问 / Qwen39 条X:Charlie Holtz38 条X:Google AI37 条X:Google DeepMind37 条X:Thariq34 条X:Peter McCrory(Anthropic 首席经济学家)31 条X:AK30 条X:百度 Baidu28 条X:AI at Meta28 条X:Boris Cherny28 条X:Deedy Das27 条X:可灵 Kling AI26 条X:Aidan Gomez(Cohere CEO)26 条X:Viggle AI26 条X:Mustafa Suleyman(Microsoft AI CEO)25 条X:Logan Kilpatrick24 条X:karminski22 条X:Josh Woodward21 条X:Noam Brown20 条X:李继刚18 条X:Andrew Milich18 条X:Odyssey18 条X:硅基流动 SiliconFlow17 条X:SpaceXAI17 条X:Tianyi Cui17 条X:DeepSeek16 条X:华为云14 条X:Sundar Pichai14 条X:智谱 Z.ai12 条X:Mark Zuckerberg12 条X:Demis Hassabis11 条X:唐杰10 条X:Ammaar Reshi10 条X:Andrew Ng(DeepLearning.AI 创始人)10 条X:Eric Mitchell10 条X:Karina Nguyen10 条X:Lee Robinson10 条X:Mistral AI10 条X:Noah Zweben10 条X:Fei-Fei Li9 条X:Jensen Huang9 条X:Kimi.ai9 条X:Barret 李靖8 条X:Jeff Dean8 条X:Michael Truell7 条X:張小珺 Xiaojùn6 条@berryxia · 历史来源430 条@vista8 · 历史来源262 条@op7418 · 历史来源231 条@sensetime_ai · 历史来源34 条@google · 历史来源30 条@getsuperintel · 历史来源9 条@latentspacepod · 历史来源9 条@android · 历史来源8 条@dreamlabla · 历史来源8 条@mannybernabe · 历史来源8 条@karpathy · 历史来源7 条@alexatallah · 历史来源6 条@ryanleeminimax · 历史来源5 条@theo · 历史来源5 条PixVerse (@PixVerse) · 历史来源5 条@aidotengineer · 历史来源4 条@dkundel · 历史来源4 条@reach_vb · 历史来源4 条@dotey · 历史来源3 条@eliebakouch · 历史来源3 条@googlechrome · 历史来源3 条@kilocode · 历史来源3 条@maxforai · 历史来源3 条@newsfromgoogle · 历史来源3 条@richardssutton · 历史来源3 条@skylermiao7 · 历史来源3 条@victorsuortiz · 历史来源3 条@ajambrosino · 历史来源2 条@akashi203 · 历史来源2 条@anatolikopadze · 历史来源2 条@andrewcurran_ · 历史来源2 条@antirez · 历史来源2 条@barrnanas · 历史来源2 条@coreyching · 历史来源2 条@deanwball · 历史来源2 条@designarena · 历史来源2 条@fal · 历史来源2 条@fellmentke · 历史来源2 条@gergelyorosz · 历史来源2 条@gmi_cloud · 历史来源2 条@gravicle · 历史来源2 条@hxiao · 历史来源2 条@jackminong · 历史来源2 条@lennysan · 历史来源2 条@mada299 · 历史来源2 条@microsoft · 历史来源2 条@mikastars39 · 历史来源2 条@mitchellh · 历史来源2 条@nabeelqu · 历史来源2 条@rudrank · 历史来源2 条@sebastienbubeck · 历史来源2 条@zan2434 · 历史来源2 条@___harald___ · 历史来源1 条@_boraturan · 历史来源1 条@0xjaniak · 历史来源1 条@0xkato · 历史来源1 条@47fucb4r8c69323 · 历史来源1 条@559hkdt · 历史来源1 条@aaliya_va · 历史来源1 条@abhikatte42 · 历史来源1 条@abhishekpatiil · 历史来源1 条@aboutberlin · 历史来源1 条@addyosmani · 历史来源1 条@agi2asi · 历史来源1 条@aiaicreate · 历史来源1 条@aimlapi · 历史来源1 条@aisaonehq · 历史来源1 条@aisystemprompt · 历史来源1 条@alemtuzlak · 历史来源1 条@alexxubyte · 历史来源1 条@alupsasca · 历史来源1 条@amasad · 历史来源1 条@ampcode · 历史来源1 条@anas_build_ · 历史来源1 条@aniketmaurya · 历史来源1 条@anitakirkovska · 历史来源1 条@anneliesgamble · 历史来源1 条@antigravity · 历史来源1 条@arafatkatze · 历史来源1 条@arena · 历史来源1 条@argofowl · 历史来源1 条@arthurmensch · 历史来源1 条@ashiknewazaj · 历史来源1 条@atabarrok · 历史来源1 条@atomic_chat_hq · 历史来源1 条@awe_automation · 历史来源1 条@awesomekling · 历史来源1 条@awscloud · 历史来源1 条@ayushagarwal · 历史来源1 条@baaadas · 历史来源1 条@bai_agi · 历史来源1 条@bbuddha_xyz · 历史来源1 条@bclavie · 历史来源1 条@beccalytics · 历史来源1 条@benfleming__ · 历史来源1 条@benhylak · 历史来源1 条@benjamineyliu · 历史来源1 条@bfl_ml · 历史来源1 条@bleysg · 历史来源1 条@bolna_dev · 历史来源1 条@bosmeny · 历史来源1 条@boxmining · 历史来源1 条@bozhou_ai · 历史来源1 条@brexhq · 历史来源1 条@brian_armstrong · 历史来源1 条@brianchew · 历史来源1 条@bridgemindai · 历史来源1 条@budgetpixel · 历史来源1 条@cahidarda · 历史来源1 条@calmpromptshq · 历史来源1 条@ce_zhang · 历史来源1 条@cedric_chee · 历史来源1 条@chaitralikakde · 历史来源1 条@chatgpt · 历史来源1 条@chatgptapp · 历史来源1 条@christinetyip · 历史来源1 条@christofsalis · 历史来源1 条@clark__labs · 历史来源1 条@cloudflaredev · 历史来源1 条@cnorth_13 · 历史来源1 条@cnzoecomeback · 历史来源1 条@cocohearts · 历史来源1 条@code_star · 历史来源1 条@codebyaurelia · 历史来源1 条@cognition · 历史来源1 条@commandcodeai · 历史来源1 条@consensusnlp · 历史来源1 条@contralabs_ai · 历史来源1 条@cozyblaze265065 · 历史来源1 条@crimedecoder · 历史来源1 条@crtr0 · 历史来源1 条@damnventures · 历史来源1 条@daniellockyer · 历史来源1 条@darioamodei · 历史来源1 条@davidmaliglowka · 历史来源1 条@davidondrej1 · 历史来源1 条@dbirker78883 · 历史来源1 条@deryatr_ · 历史来源1 条@devfun · 历史来源1 条@diegocabezas01 · 历史来源1 条@digitalocean · 历史来源1 条@dimillian · 历史来源1 条@dimitrispapail · 历史来源1 条@discussingfilm · 历史来源1 条@dkthomp · 历史来源1 条@dmitryrybin1 · 历史来源1 条@dmsobol · 历史来源1 条@douglasyaody · 历史来源1 条@duckduckgo · 历史来源1 条@easyrouterio · 历史来源1 条@edgardobriban · 历史来源1 条@eisokant · 历史来源1 条@elliotarledge · 历史来源1 条@encrypted · 历史来源1 条@endpointarena · 历史来源1 条@envato · 历史来源1 条@escanorreloaded · 历史来源1 条@esrtweet · 历史来源1 条@ethanhe_42 · 历史来源1 条@eu_commission · 历史来源1 条@fba · 历史来源1 条@fdavidsont · 历史来源1 条@figmaweave · 历史来源1 条@finn_meeks · 历史来源1 条@first_tree_ai · 历史来源1 条@flavioad · 历史来源1 条@flowith · 历史来源1 条@fminzhou · 历史来源1 条@freddie_spirit · 历史来源1 条@frydwia · 历史来源1 条@futurestacked · 历史来源1 条@garrettlord · 历史来源1 条@garrytan · 历史来源1 条@gavinsbaker · 历史来源1 条@GayaniFigma · 历史来源1 条@genspark_ai · 历史来源1 条@gitlawb · 历史来源1 条@gneubig · 历史来源1 条@gokulr · 历史来源1 条@goodfireai · 历史来源1 条@goodnesmbakara · 历史来源1 条@googleaistudio · 历史来源1 条@gordic_aleksa · 历史来源1 条@gro_tsen · 历史来源1 条@hangsiin · 历史来源1 条@happycapyai · 历史来源1 条@haydenbleasel · 历史来源1 条@helloiamleonie · 历史来源1 条@hey_asiif · 历史来源1 条@hilbertspaess · 历史来源1 条@howtoprompt__ · 历史来源1 条@hq4ai · 历史来源1 条@hypersoren · 历史来源1 条@ianbremmer · 历史来源1 条@id_aa_carmack · 历史来源1 条@interaction · 历史来源1 条@intology · 历史来源1 条@iron_redux · 历史来源1 条@ithilgore · 历史来源1 条@itsreallyvivek · 历史来源1 条@jamesjyu · 历史来源1 条@jameszmsun · 历史来源1 条@jason_young1231 · 历史来源1 条@jawad_rahman_ · 历史来源1 条@jaydendavisnc · 历史来源1 条@jeffbarg · 历史来源1 条@jenzhuscott · 历史来源1 条@jiayuan_jy · 历史来源1 条@jilles · 历史来源1 条@jimcramer · 历史来源1 条@jimsyoung_ · 历史来源1 条@jinjingliang · 历史来源1 条@jjacky · 历史来源1 条@jjackyliang · 历史来源1 条@joefioti · 历史来源1 条@joi___ai · 历史来源1 条@joinhandshake · 历史来源1 条@joinpursuit · 历史来源1 条@joulee · 历史来源1 条@jsconfasia · 历史来源1 条@jsrailton · 历史来源1 条@juminoz · 历史来源1 条@kaizero_ainta · 历史来源1 条@karanganesan · 历史来源1 条@kdaigle · 历史来源1 条@kentherogers · 历史来源1 条@kevinsays · 历史来源1 条@khudonogov · 历史来源1 条@koraykv · 历史来源1 条@kotekjedi_ml · 历史来源1 条@kuberwastaken · 历史来源1 条@kurz_gesagt · 历史来源1 条@kwindla · 历史来源1 条@lafalcemateo · 历史来源1 条@lakshyaaagrawal · 历史来源1 条@larrylv · 历史来源1 条@layoffai · 历史来源1 条@levinstanley · 历史来源1 条@lifeofjer · 历史来源1 条@livekit · 历史来源1 条@lostinlatencyx · 历史来源1 条@lotte_verheyden · 历史来源1 条@lqiao · 历史来源1 条@luciushq · 历史来源1 条@luckeyfaraday · 历史来源1 条@lukaspet · 历史来源1 条@madhavsinghal_ · 历史来源1 条@manassharmahere · 历史来源1 条@markiewagner · 历史来源1 条@marksaroufim · 历史来源1 条@marsxiang_ · 历史来源1 条@maseehg_ · 历史来源1 条@mattshumer_ · 历史来源1 条@mem0ai · 历史来源1 条@mengto · 历史来源1 条@merettm · 历史来源1 条@micahcarroll · 历史来源1 条@michael_chomsky · 历史来源1 条@michaelarnaldi · 历史来源1 条@microsoftai · 历史来源1 条@mike_acton · 历史来源1 条@mikeyyyzhao · 历史来源1 条@minchoi · 历史来源1 条@minimaxagent · 历史来源1 条@minu_who · 历史来源1 条@mkbhd · 历史来源1 条@modal · 历史来源1 条@moritzthuening · 历史来源1 条@moxie · 历史来源1 条@mstockton · 历史来源1 条@mtslive · 历史来源1 条@multimodalart · 历史来源1 条@neelnanda5 · 历史来源1 条@neilrahilly · 历史来源1 条@nickbaumann_ · 历史来源1 条@nirantk · 历史来源1 条@noemititarenco · 历史来源1 条@notjazii · 历史来源1 条@nousresearch · 历史来源1 条@oblomovius · 历史来源1 条@ollama · 历史来源1 条@onlyterp · 历史来源1 条@onlyzhynx · 历史来源1 条@organicgpt · 历史来源1 条@orgrem · 历史来源1 条@p0 · 历史来源1 条@palantirtech · 历史来源1 条@palmerluckey · 历史来源1 条@pandatalk8 · 历史来源1 条@parishilton · 历史来源1 条@patrickcarlyle · 历史来源1 条@patricktoulme · 历史来源1 条@paulg · 历史来源1 条@paulsolt · 历史来源1 条@pbdtokenrouter · 历史来源1 条@pererabinoy · 历史来源1 条@philhchen · 历史来源1 条@pirroh · 历史来源1 条@pjaccetturo · 历史来源1 条@postlive · 历史来源1 条@pranaveight · 历史来源1 条@prathamdby · 历史来源1 条@prince_canuma · 历史来源1 条@pumpkherm · 历史来源1 条@pvncher · 历史来源1 条@qiaoqiao2001 · 历史来源1 条@rajveerbach · 历史来源1 条@randyhaddad6 · 历史来源1 条@rauchg · 历史来源1 条@raveeshbhalla · 历史来源1 条@rayanpal_ · 历史来源1 条@rayfernando1337 · 历史来源1 条@redpoint · 历史来源1 条@ric_rtp · 历史来源1 条@richardsocher · 历史来源1 条@rileybrown · 历史来源1 条@robertvaradan · 历史来源1 条@ronshepherd · 历史来源1 条@rosmine · 历史来源1 条@rthiago · 历史来源1 条@ruben_kostard · 历史来源1 条@runware · 历史来源1 条@rvivek · 历史来源1 条@ryanjunejo · 历史来源1 条@safaricheung · 历史来源1 条@samuelstroschei · 历史来源1 条@sanmking · 历史来源1 条@saranormous · 历史来源1 条@savinovnikolay · 历史来源1 条@scale_ai · 历史来源1 条@scaling01 · 历史来源1 条@sdaily_ai · 历史来源1 条@secscottbessent · 历史来源1 条@seltaa_ · 历史来源1 条@sergiopaniego · 历史来源1 条@servasyy_ai · 历史来源1 条@sethltx · 历史来源1 条@shashankgoyal95 · 历史来源1 条@sherryyanjiang · 历史来源1 条@sherylhsu02 · 历史来源1 条@shl · 历史来源1 条@sighjith · 历史来源1 条@simistern · 历史来源1 条@southpkcommons · 历史来源1 条@sriramkri · 历史来源1 条@sshoaibali · 历史来源1 条@stalkermustang · 历史来源1 条@status_effects · 历史来源1 条@stevencheng · 历史来源1 条@stockanalystpro · 历史来源1 条@suekhim · 历史来源1 条@sultanalfardan · 历史来源1 条@suraj_sharma14 · 历史来源1 条@swisscheese4299 · 历史来源1 条@swmansion · 历史来源1 条@systematicls · 历史来源1 条@teksedge · 历史来源1 条@tftc21 · 历史来源1 条@theahmadosman · 历史来源1 条@themidasproj · 历史来源1 条@theonejvo · 历史来源1 条@therealadamg · 历史来源1 条@timsoulo · 历史来源1 条@tmuxvim · 历史来源1 条@tobi · 历史来源1 条@togethercompute · 历史来源1 条@trackernetwork · 历史来源1 条@trustkerneltech · 历史来源1 条@ttunguz · 历史来源1 条@tuhinchakr · 历史来源1 条@twistartups · 历史来源1 条@ubermenscchh · 历史来源1 条@udayan_w · 历史来源1 条@usefastlane · 历史来源1 条@uzyn · 历史来源1 条@valeriocapraro · 历史来源1 条@vasuman · 历史来源1 条@vdbergrianne · 历史来源1 条@vibeguessing · 历史来源1 条@victoriakimse · 历史来源1 条@victoriawu77 · 历史来源1 条@victortaelin · 历史来源1 条@vikaskansalhq · 历史来源1 条@volchika · 历史来源1 条@walden_yan · 历史来源1 条@warpdotdev · 历史来源1 条@waynesutton · 历史来源1 条@wesroth · 历史来源1 条@whosamberella · 历史来源1 条@xdinodeer · 历史来源1 条@xicilion · 历史来源1 条@xucian_ · 历史来源1 条@yacinemtb · 历史来源1 条@yaojingang · 历史来源1 条@yoheinakajima · 历史来源1 条@yongquanyq · 历史来源1 条@youtubejocoding · 历史来源1 条@yusufg · 历史来源1 条@zachbussey · 历史来源1 条@zeddotdev · 历史来源1 条@zeroxkyle · 历史来源1 条@zhenthebuilder · 历史来源1 条@zicohacks · 历史来源1 条@zixuanli_ · 历史来源1 条@zymazza · 历史来源1 条
312 条AI 相关新闻 · 最新在前
今天10月6日周二
  1. elvis42

    和 Jev 一样,我相信 RL 会解锁更多类似这样的成果,大幅削减关键智能体操作的成本。 flow-1 在性能上很有竞争力,但在发现智能体追踪中的故障方面能大幅节省成本。 RSI 不只适用于通用智能。它同样会加速专用智能。

    引用Robert@skull8888888888

    Introducing flow-1, our new model trained with RL to find errors in agent traces. It matches GPT-6-sol in trace intelligence while being 23x cheaper. It also costs 25% less to run than GPT-6-luna. flow-1 finally makes it possible to monitor and understand every agent run, without sampling. 1/6

  2. elvis56

    Eon 推出 Era,可生成一个完整的模拟企业环境,覆盖 Salesforce、Slack、Jira、Zendesk、Gong、Deel 等系统及云数据库和存储,Agent 通过 MCP 和 API 接口与之交互。每个系统模拟真实厂商行为,包括 rate limits、分页和错误码;因环境由 Era 生成,平台掌握精确的 ground truth,可重置后对同一任务重复运行,以测量改动模型或提示词后的影响。产品今日上线且免费,与 NVIDIA、Decart、Composio 等有研究合作。作者 Elvis Saravia 评价称模拟企业对改进智能体产品很重要,因为好的 Agent 评测需要行为接近真实公司的环境。

    引用Ofir Ehrlich@OfirEhrlich

    You can’t safely test an enterprise agent on a real company’s data. So we built a company for it to work in. Era is live today, and it’s free. It generates a complete simulated enterprise that behaves like a real one across Salesforce, Slack, Jira, Zendesk, Gong, Deel and more, along with cloud databases and storage. Agents interact with it through live MCP and API interfaces. People leave. Deals change. Records get duplicated. Permissions differ across systems. And because Era generated the company, it knows the exact ground truth. Test, benchmark and improve agents against realistic enterprise workloads, use the failures for targeted post-training, then rerun the same environment to measure the impact. Huge thanks to our research partners @NVIDIA, @Decart, @Composio, @openlayerco, @Deel, @Eragon and @Plurai, with more coming soon.

  3. elvis45

    推荐。LLM 智能体喜欢结构,所以语料库能改善智能体搜索并不意外。

    引用DAIR.AI@dair_ai

    Banger paper from Microsoft and colleagues. If you run agents that search a large document collection, this one is worth your time. (bookmark it) They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it. The original documents stay in place. The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again. Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%. It also beats an LLM Wiki layer and three other navigation layers. The map can be built without LLM calls and updated as new documents arrive. Paper: https://arxiv.org/abs/2609.37226 Chat with Paper: https://academy.dair.ai/papers/follow-the-entities-a-corpus-map-for-agentic-search-2609.37226

10月5日周一
  1. elvis64

    Nolla Health 成为美国首个获得监管批准、可用 AI 开具初始处方的机构,作者 Elvis Saravia 转发了这一消息。AI 目前可在犹他州用于痤疮治疗,覆盖接诊、皮肤扫描、评估、处方和随访,必要时由临床医生介入;引用内容称其可能是全球首个端到端 AI 医生。

    引用Luis Wenus@luiswenus

    Today, Nolla Health became the first organization in the U.S. (and possibly the world) to receive regulatory approval for an AI to issue initial prescriptions. This makes Nolla the first ever actual end-to-end AI doctor.

  2. elvis45

    Elvis Saravia 主张个人智能体应当完全自建自持,从 traces 到记忆再到所有输出都归自己所有,而不是等待厂商修复 bug 或解决安全顾虑。他认为 AI 让构建和维护智能体变得简单,可从零开始保持极简并随需求演进,也可基于开源方案 Pi Durable 起步。

    引用Kun Chen@kunchenguid

    there’s something quite awkward about all the personal agents in the current hype cycle muse, grok bot, instinct, dots, and whatever google, anthropic will come up with none of them is “mine” i’d be trusting a vendor for some of the most sensitive data about me and bet on them handling it with extreme responsibility. what if they have a data leakage incident? what if they hired a wrong employee? what if they simply break bad? i’d also be betting on their model. right now i’ve set up a lot of my stuff in grok bot, but what if their model falls behind? what if another model becomes 10x better on pure intelligence? then i’ll either miss out on a better assistant, or have to bite the bullet and do a full migration i’d also have to bet on their product. they may not build the features i want. they may not allow me to customize enough. and one day my assistant may show me ads that feels like way too much betting than what i’d be comfortable with, if i truly want literally everything in my life to go through and be taken care of by an assistant with all that considered, i’m more bullish on open source personal agents that run on people’s own computers. but i think these open source projects have to break out of the assumption that their users are developers and they can just throw a repo at them and ask them to launch a terminal a well polished, fool proof, community maintained, vendor agnostic, free personal agent that everyone can run by themselves without locking into a SaaS - that’s what i think many people will need

  3. elvis45

    如今,我花在验证智能体写的代码上的时间,比写代码还多。 这是所谓软件工厂的一大瓶颈。 来自 @ContextQa 的 Ship 会接收 Slack 或 Linear 中报告的 bug,复现它,并把上下文交给 Claude Code 或 Codex 去修复。 编程智能体需要这种反馈循环,才能更擅长修复自己的 bug。

    引用Deep Barot@deepcabinwala

    Everyone got a coding agent. Nobody got a QA engineer. Until today. Meet Ship, your Autonomous Quality Engineer. It tests deployed PRs, reproduces bugs from Slack and Linear, and hands Claude or Codex the context to fix them. Try it free: https://ship.contextqa.com

  4. elvis59

    Elvis Saravia 在引用 Karpathy 关于用 ASD-STE100 写作、图表、网页和讲解视频等方式理解 LLM 输出的帖子时,分享了自己数月来实验的通用人机协作界面。

    引用Andrej Karpathy@karpathy

    We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.

10月4日周日
  1. elvis32

    长期使用终端的开发者表示,AI 智能体已让他完全放弃终端操作,如今可以委托给智能体的工作大幅增加,且已有明显更好的智能体交互界面。他认为终端时代积累的系统思维仍是 AI 时代的护城河,有助于设计实用的沙箱及其上层抽象。

    引用Yuchen Jin@Yuchenj_UW

    I was a terminal person for 15+ years. I loved Vim, knew all the shortcuts, and a black terminal made me feel like I was a cool hacker in The Matrix. But terminals assume humans operate computers directly via files, commands, processes. Agents changed all that. You just say your intent. The agent operates the machine, using the same programming languages and Linux commands we spent years learning. I’m still glad I learned computer systems before the AI era though. Understanding what sits beneath the abstraction makes you a better systems thinker. That’s a moat in the AI era.

  2. elvis32

    Elvis Saravia 认为通过 CLI 使用编码智能体的方式已经过时,交互正直接发生在智能体之间。他的做法是混合使用高层持久化智能体(如 dots)和专用智能体(如 codex desktop 会话),由持久化智能体管理专用智能体,并借助自建的智能体编排器获得灵活性。

    引用Yuchen Jin@Yuchenj_UW

    I haven’t touched Claude Code or Codex CLI in a while. The terminal era is over imo. It's the wrong interface for coding agents. Tabs are ephemeral, but context is persistent, and managing 30 tabs is pure cognitive overhead. I don’t really need an IDE like Cursor either. I rarely navigate the whole codebase anymore. The new primitive is the agent, not the file. (Codex desktop app is the best agentic UI for now. But we’re still early.)

  3. elvis59

    Meta Superintelligence Labs 论文发现,预训练 LLM 配上轻量推理框架后可在足够测试时预算下超越 RL 后训练版本。后训练把任务推向必解或必不解,提升一致性但降低覆盖,作者称之为 Sharpening Tax;在 14 类基座/后训练配对、42 组设置中普遍存在,其 PTGS 方法按提示词难度自适应调整 RL 采样温度,付出更小代价并提升 pass@1。论文见 https://academy.dair.ai/papers/sharpening-tax-in-post-training-2610.01509。

    引用DAIR.AI@dair_ai

    Banger paper from Meta Superintelligence Labs. They find something super interesting and unexpected. (bookmark it) Base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples. Post-trained models win on pass@1. At large K, base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop. This is because post-training pushes each task toward always solved or never solved. Consistency goes up, and coverage goes down. The authors call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it shows up in most settings, grows with model size, and can be estimated from a few rollouts. Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass@1. Paper: https://academy.dair.ai/papers/sharpening-tax-in-post-training-2610.01509

10月3日周六
  1. elvis38

    Kled V3 可将实验室指定的数据采集任务在 72 小时内分发至其 50 万+ 自愿贡献者网络,覆盖图像、视频、音频、文本与标注共 108 个可配置模板。贡献者用手机按说明和合格示例采集真实世界数据,形成"定位模型失败→生成采集任务→获取新样本→再训练评估"的更紧反馈闭环。

    引用Kled AI@useKled

    This is Kled V3. We've solved data collection for artificial general intelligence. Any consumer dataset that can exist, can now be collected from physical reality in under 72 hours. All powered by the largest and most comprehensive data application layer on the planet. (Thread)

10月2日周五
  1. elvis48

    AgentWorld 将 3 到 20 个不同角色的 LLM 智能体放入游戏沙盒,执行 50+ 轮的长程任务,智能体无法看到彼此内部状态,只能通过消息和共享计划协调。Gemini 3 Flash 任务成功率最高,为 52.0%;协调类任务最难,成功率仅 12%,常见失败包括沟通中断、角色混淆和共享计划丢失。论文显示,多智能体团队中不到三分之一的行动真正有助于完成任务。

  2. elvis50

    我认为更快的推理是编程智能体下一个重大突破之一。 Volantis 正在利用光学技术为每颗芯片提供远超以往的内存和更高的内存带宽。他们的目标是在超过 10T 参数的模型上,实现每用户每秒高达 10,000 tokens。这太疯狂了! 以那种速度,今天需要数小时的编程智能体可以在几分钟内完成。 绝对是我最近见过的最令人兴奋的融资之一。

    引用Tapa Ghosh@semiDL

    Excited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: https://volantissemi.ai/news-insights/our-88m-series-a-demolishing-the-memory-wall-with-photonics-post

  3. elvis48

    我觉得 AI 语音缺了一个质检层。 在我构建的大多数应用里,语音听起来已经很棒了。但真正让我在生产环境翻车的,是模型把“Q3”、“v1.2”或某个品牌名读错。 这对人们对语音智能体的感知影响巨大。 @OnepinAI 正在解决这个问题。在你听到之前,他们会检查每一行的自然度、清晰度和发音。 当某个词读错时,它会直接修复那个词,而不是重新生成整行。

    引用Soohyun Bae@RealSoohyunBae

    Your AI voice sounds human. So why can't it say your product's name? A great AI voice reads "Porsche Taycan" as TAY-can. Porsche says TIE-kahn. It guessed from the spelling, and nobody caught it, because nobody listens to line 1,200. Today we're launching Onepin: the production step after text-to-speech. It checks every line of voiceover before it ships using the voices you already work with. Onepin can: ➤ Check people's and product names against a 4-million-word pronunciation dictionary ➤ Spell out prices and dates before the voice speaks ➤ Score every line of audio for naturalness, clarity and word accuracy ➤ Fix the one wrong word in the same voice, without re-rendering the take Works with your voice subscription on @ElevenLabs, @OpenAI, @Google and 30+ more. No phonetic spellings to type. No re-rolls. No switching providers. Free to start, no credit card required. Hear the before and after in the thread ⬇️

  4. elvis48

    webAI 开源 3.6B 参数模型 TwIL-LM3-Pro,可在普通电脑本地运行,BIG-Bench Hard 得分 95.4,远超 Qwen3-8B 的 63.7。其训练配方为:形式逻辑微调后将权重合并回基座模型,再用程序化验证器做 RL,逻辑分数提升且通用推理保持稳定。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  5. elvis58

    Tavus 发布 Griffin,称其为首个通过视频图灵测试的 Human Interaction Model,48% 的实时对话者认为它是真人,此前系统通过率低于 3%。它是一个 video-to-video 模型,能边看边听、被打断和打断对方,并 reacting 房间内发生的情况;在 NVIDIA Video Full-Duplex Benchmark 上生成得分 3.83/5,真人为 3.92。

    引用Tavus@tavus

    Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).

  6. elvis57

    Elvis Saravia 转发 Arceus Legal 融资消息并提出观点,预计更多 AI 公司会自己运营服务并出售结果。Arceus Legal 宣布获得 1700 万美元融资,由 greycroftvc 领投;客户通过 Slack 发送合同,AI 收集上下文,执业律师审查每份工作产出,合同审查平均 3 到 5 小时完成,按开工前约定的固定费用收费。

    引用Mac Liu@themacliu

    I’m excited to announce that @arceuslegal is launching with $17M in funding, led by @greycroftvc, with participation from @craft_ventures, @spc, and others. As a founder, I always hated how helpless I felt working with law firms. I went through four or five different firms and somehow the experience was always the same. I’d be waiting on something important to our business with no idea when I’d hear back. I’d have to re-explain our business over and over again. And I dreaded jumping on calls because I knew every minute was costing me money. We started Arceus because we believe every business deserves a better law firm. One that moves faster, costs less, and puts the client first. And we’re just getting started. ↓

10月1日周四
  1. elvis33

    自己动手搭建 harness 吧,各位! 搭建我自己的 meta harness 最大的好处之一,就是切换模型非常顺畅。 随着模型发布周期不断缩短,我觉得这一点正变得越来越重要。RSI 只会让发布节奏进一步加快。 这是自去年 11 月以来我做过的最好的决定之一。 切换模型确实可能搞崩很多东西,但它不该如此。如果会崩,说明你的配置不适合我们当下这个疯狂的发布节奏。你现在就得解决这个问题。

9月13日周日
  1. @omarsar026

    Elvis Saravia 认为,大量 YC 创业者正投入构建领域专用 harness,因为深耕某一领域后会发现,harness 是智能体时代保持竞争力的关键。从产品看,harness 能开辟新的交互界面与体验;从技术看,它是维护用户与产品交互框架和最佳实践的载体。他建议新手把一份 harness 论文清单交给自己的智能体学习,已有经验的开发者则可为自己的领域动手构建一个。

  2. @omarsar041

    一项新研究提出循环架构与配套训练方法,在不增加参数的前提下提升模型推理能力,在六个推理基准中的五个上超越此前的循环模型。该方法用局部去噪目标训练循环过程,逐步降低噪声并共享同一噪声样本,使每次更新与下一次更新关联起来。推理时模型沿概率流运行,更细的时间网格可增加计算量,不同起始噪声在存在多解的任务上可产生不同有效答案。

9月12日周六
  1. @omarsar029

    Elvis Saravia 建议对输入 API 或聊天模型的数据保持高度警惕,认为当前缺乏保障,尤其在前沿实验室推进 RSI 的背景下。他因此看好开源与开放权重模型,并强调企业未来价值来自客户、用户、员工及复杂系统交互中涌现的智能,分享这些轨迹等于送出智能栈的一部分。他表示并不反对闭源模型,只是更谨慎地路由任务,自建 harness 即可做到。

  2. @omarsar044

    HKUST 提出 AgentZip,针对并行运行大量智能体沙箱时的高内存冗余,将沙箱内存压缩最高 8.7 倍(Linux 配置为 2.1 倍)。研究者测得 76% 至 96% 的页面存在模板相对或跨沙箱冗余,AgentZip 同时针对模板和同级沙箱压缩页面,包括相似但不完全相同的页面。它在智能体等待 LLM 时执行压缩,并在恢复时预取页面,将激进压缩带来的 3.1 倍减速降至 1.40 倍。