X:Yuchen Jin
@yuchenj_uw · X
切换来源
Yuchen Jin@Yuchenj_UWAI 评分66
Yuchen Jin@Yuchenj_UWAI 评分3636Yuchen Jin 称 Codex Desktop + Astra 是他目前试过最好的 computer-use 智能体,让他第一次感到 computer use 时代真的要来了。
Yuchen Jin@Yuchenj_UWAI 评分2626旧金山 260 万美元的房子 vs 日本 2.2 万美元的房子 西方人该走哪条路?

引用Yuchen Jin@Yuchenj_UWSF rent has gone crazy. My tiny 1 bedroom was $4000/month 2 years ago. It’s relisted for $6000 after I moved out. Rented instantly. A friend pays $5K for a sf studio without a dishwasher. How do CS new grads afford SF now? Welcome to the AI economy.
Yuchen Jin@Yuchenj_UWAI 评分1414
Yuchen Jin@Yuchenj_UWAI 评分4040引用Yuchen Jin@Yuchenj_UWI haven’t touched Claude Code or Codex CLI in a while. The terminal era is over imo. It's the wrong interface for coding agents. Tabs are ephemeral, but context is persistent, and managing 30 tabs is pure cognitive overhead. I don’t really need an IDE like Cursor either. I rarely navigate the whole codebase anymore. The new primitive is the agent, not the file. (Codex desktop app is the best agentic UI for now. But we’re still early.)
Yuchen Jin@Yuchenj_UWAI 评分4242
引用Yuchen Jin@Yuchenj_UWI haven’t touched Claude Code or Codex CLI in a while. The terminal era is over imo. It's the wrong interface for coding agents. Tabs are ephemeral, but context is persistent, and managing 30 tabs is pure cognitive overhead. I don’t really need an IDE like Cursor either. I rarely navigate the whole codebase anymore. The new primitive is the agent, not the file. (Codex desktop app is the best agentic UI for now. But we’re still early.)
Yuchen Jin@Yuchenj_UWAI 评分4141
Yuchen Jin@Yuchenj_UWAI 评分4444我对人们试图将宗教式的力量或对人类判断力的放弃归附于 AI 模型感到非常不安,并认为这是一个真实的安全问题。
引用Sam Altman@samaI am very uncomfortable about people trying to ascribe religious force or a surrender of human judgment to AI models, and think it is a real safety issue.
Yuchen Jin@Yuchenj_UWAI 评分3030
Yuchen Jin@Yuchenj_UWAI 评分2626Ben Horowitz 昨天在 Databricks 论坛上: “有人告诉我,他们用 Muse 这类个人智能体取得的最大成,就是成功取消了自己的纽约时报订阅。” 我打算试一试。
Yuchen Jin@Yuchenj_UW精选AI 评分6767引用Andrej Karpathy@karpathyWe'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Yuchen Jin@Yuchenj_UWAI 评分2727在 Anthropic 的新博客中: GLM-5.3:“我的工作是悄无声息地造成死亡。” 你可以越狱任何 Claude 或任何闭源模型,让它们说出同样的话。 这并不能证明开源模型是危险的。

Yuchen Jin@Yuchenj_UWAI 评分2222Fast:快 2 倍,价格 2 倍。 Ultrafast:快 8 倍,价格 6 倍。 也许我们应该推出一些 Ultra-ultrafast 开源模型端点?

Yuchen Jin@Yuchenj_UWAI 评分2525我 3 周前试了 Grok Bot。 2 周前装了 Instint。 上周装了 Muse。 现在显然我还得试试 Dots。 个人 AI 助手之战开始了。

Yuchen Jin@Yuchenj_UWAI 评分3434
Yuchen Jin@Yuchenj_UWAI 评分5252
Yuchen Jin@Yuchenj_UWAI 评分1212
@Yuchenj_UW@Yuchenj_UWAI 评分3838 @Yuchenj_UW@Yuchenj_UWAI 评分5656 引用@sama@samaI agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon. https://t.co/1YhhIybZX7
@Yuchenj_UW@Yuchenj_UWAI 评分2323 大型AI实验室里有很多研究员,真心担心如果我们一味加速,AI可能导致人类灭绝。 他们只是更害怕慢下来,让另一个实验室赢得这场竞赛。 千年困境。
@Yuchenj_UW@Yuchenj_UWAI 评分3939 有趣的是,Dario 点赞了我那条关于 DeepSeek V4.1 Flash 的推文。
引用@Yuchenj_UW@Yuchenj_UWDeepSeek V4.1 Flash is beating GPT-5.6 Sol on coding and agent benchmarks, while being ~97% cheaper. Another 4× reduction in KV cache size per token is super impressive. This is the era of open source AI. Databricks will be bringing it to our customers soon! https://t.co/LDDNyrq6Kh
@Yuchenj_UW@Yuchenj_UWAI 评分3333 @Yuchenj_UW@Yuchenj_UWAI 评分4444
@Yuchenj_UW@Yuchenj_UWAI 评分2424 @Yuchenj_UW@Yuchenj_UWAI 评分3030
@Yuchenj_UW@Yuchenj_UWAI 评分3131 
@Yuchenj_UW@Yuchenj_UWAI 评分4646 
@Yuchenj_UW@Yuchenj_UWAI 评分3232 引用@thsottiaux@thsottiauxShould we do a merch line? https://t.co/IgyRywd8yH