AI 导读
小北部分认同 Karpathy 关于 AI 输出将从 markdown 演进到 HTML、最终走向扩散模型直出交互视频的判断,认为 HTML 在仪表盘、对比和小交互上确实是质变。
正文
Karpathy 说视觉是 AI 输出的首选通道,所以未来 HTML 会取代 markdown,再往后是神经视频。
一半同意吧,HTML 在做仪表盘、做对比和一些小交互这类东西上确实是质变,markdown 给不了。
但视觉是首选输出,这个说的有点太满了。
看文字本身就是视觉处理啊,不是只有图形界面才用上眼睛。
并且带宽不等于高效,视觉皮层是宽,但读文本走的是高度优化过的符号通路,未必比解析复杂的布局慢。
一些代码、推理过程,还有需要精确表达的东西,纯文本反而最舒服。HTML 是有隐性成本的,很重也很难二次编辑。
至于终点是扩散模型直出交互视频,技术上不是科幻。
但我有点怀疑它该不该成为通用输出形态,可交互神经世界作为体验是加分,作为默认 I/O 可能丢的比换来的多。
This works really well btw, at the end of your query ask your LLM to "structure your response as HTML", then view the generated file in your browser. I've also had some success asking the LLM to present its output as slideshows, etc. More generally, imo audio is the human-preferred input to AIs but vision (images/animations/video) is the preferred output from them. Around a ~third of our brains are a massively parallel processor dedicated to vision, it is the 10-lane superhighway of information into brain. As AI improves, I think we'll see a progression that takes advantage: 1) raw text (hard/effortful to read) 2) markdown (bold, italic, headings, tables, a bit easier on the eyes) <-- current default 3) HTML (still procedural with underlying code, but a lot more flexibility on the graphics, layout, even interactivity) <-- early but forming new good default ...4,5,6,... n) interactive neural videos/simulations Imo the extrapolation (though the technology doesn't exist just yet) ends in some kind of interactive videos generated directly by a diffusion neural net. Many open questions as to how exact/procedural "Software 1.0" artifacts (e.g. interactive simulations) may be woven together with neural artifacts (diffusion grids), but generally something in the direction of the recently viral nitter.net/zan2434/status/2046982… There are also improvements necessary and pending at the input. Audio nor text nor video alone are not enough, e.g. I feel a need to point/gesture to things on the screen, similar to all the things you would do with a person physically next to you and your computer screen. TLDR The input/output mind meld between humans and AIs is ongoing and there is a lot of work to do and significant progress to be made, way before jumping all the way into neuralink-esque BCIs and all that. For what's worth exploring at the current stage, hot tip try ask for HTML.在 X 查看被引用的帖子
来源:@frxiaobei · x.com