X:Ethan Mollick
@emollick · X
切换来源
@emollick@emollickAI 评分2121 @emollick@emollickAI 评分44 这些回复糟糕得离谱,我觉得所有回复这条的机器人都应该至少先试着玩玩这些游戏。(《Masque of the Red Death》引人入胜且节奏快)
@emollick@emollickAI 评分2525 

@emollick@emollickAI 评分1616 它认真对待了生成独特游戏的要求,每个游戏都用了不同的方法。 开源,欢迎随意修改:https://t.co/E2BgsFhicC https://t.co/mMSWYwmaMF

@emollick@emollickAI 评分3434 
@emollick@emollickAI 评分2424 AI 奖励专业能力(至少目前如此)。专业能力让你能判断 AI 输出的质量,并快速找到锯齿状前沿的形状。它还让你有更多选择去尝试提升质量,因为你知道该要求哪些改动。非专家往往只能困于默认设置。
@emollick@emollickAI 评分2727 @emollick@emollickAI 评分1717 @emollick@emollickAI 评分2020 

@emollick@emollickAI 评分1616 
@emollick@emollickAI 评分2020 注意这与隐私是两回事——你可以在 OpenAI 和 Claude 中关闭该选项,确保你的数据不被用于训练。你也可以使用开放权重模型,但随着它们不断改进,它们也会有同样的问题。
@emollick@emollickAI 评分2626 @emollick@emollickAI 评分2323 @emollick@emollickAI 评分2727 @emollick@emollickAI 评分1111 我确实把游戏玩到了 80%,看起来全部 60 个房间都已完整构建,所有谜题也都可解。如果你发现任何问题,我会修复。https://t.co/vXPuQezftS

@emollick@emollickAI 评分1212 最初的画作(从引用的推文可以看出,一年前图像生成器能画出这栋建筑的图像就已经很惊人了,勉强算是)。https://t.co/DHBx41HMHK


@emollick@emollickAI 评分3232 
@emollick@emollickAI 评分2222 这比一般的 vibe-coded 游戏更成功,原因之一是它以 Zork 为基础,填补了 AI 最薄弱环节的参差前沿(连贯的长线叙事、谜题设计、写作),同时把 AI 用在它最擅长的地方。
@emollick@emollickAI 评分1818 在和很多人聊 AI 的过程中,我学到的一件事是:他们既可能担心 AI 带来的影响,又对自己使用 AI 感到非常兴奋。我觉得这个平台上的人往往把许多人对 AI 的态度看得比实际简单得多。
@emollick@emollickAI 评分77 好吧。(Astra 是用 VBA 做的这个)https://t.co/uvc1V3oijr https://t.co/aCQvziTZrZ

@emollick@emollickAI 评分4242 引用@ArtificialAnlys@ArtificialAnlysAnnouncing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index v4.2 changelog: + AA-Briefcase, our agentic knowledge work evaluation with a private test set + @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! Intelligence Index v4.2 changes in detail: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5. ➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure. Key results: ➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google ➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier ➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
@emollick@emollickAI 评分2727 

@emollick@emollickAI 评分3333 
@emollick@emollickAI 评分3434 
@emollick@emollickAI 评分1313 让 GPT-5.6 Astra 在搭建一个很酷的东西,应该很快就完成了。懂的都懂(不懂的话,反正我很快也会把整个东西发出来)。https://t.co/EJMWF9Wt6w

@emollick@emollickAI 评分2121 有趣的是,费马大定理的证明描述虽然简短,却仍然带着浓浓的 Claude 味道(“为每一步命名,并标注承载它的 Lean 定理”)。https://t.co/VYbKX28QeL
@emollick@emollickAI 评分2525 @emollick@emollickAI 评分3232 

@emollick@emollickAI 评分3030 一方面,我能在不到几个小时里就得到准确、未经 p hacking 的原创研究论文,这绝对令人惊叹。另一方面,这些结果不算垃圾,只是不够有趣。对 AI 来说,即便有提示词,解决研究品味也是个难题。
@emollick@emollickAI 评分2727 



@emollick@emollickAI 评分3535
引用@emollick@emollickThe drowned neo-gothic tower twigl shader created by Fable 5.1 with the same prompt. (compare to Fable 5 in the quoted tweet, and other models before that) https://t.co/b3OCbBc9g6 https://t.co/6OWcg9Rbcj
@emollick@emollickAI 评分22 @emollick@emollickAI 评分1717 @emollick@emollickAI 评分66 @emollick@emollickAI 评分1919 @emollick@emollickAI 评分3535 
@emollick@emollickAI 评分3131 @emollick@emollickAI 评分2323 如果你还没试过,这里有全程解说导览、历史链接,你可以阅读卷轴,还可以快进到各个场景,以及关于图书馆衰败及其多次火灾的各种理论等等。 开源地址:https://t.co/39tMSzkLU1
@emollick@emollickAI 评分4444 
@emollick@emollickAI 评分2121 Astra 一个最微妙且有价值的优点是它保持了更好的心智理论:在最终产物中,你几乎不会看到对之前草稿或构建过程中所做工作的奇怪引用,而且它在长时间运行时漂移更少。不完美,但非常好。