跳到正文

全部动态

今日 38 条
9月28日周一
  1. Hugging Face Blog75

    H Company 发布 Holo4 系列通用计算机操作智能体模型

    H Company 发布 Holo4 智能体模型系列,包含 27B dense 和 35B-A3B MoE 两个尺寸,并附带基于 Nemotron 3 Nano Omni 后训练的 Holotron4 Nano。

    推荐理由:官方发布给出了跨 GUI、代码、MCP 和 API 四类接口的统一智能体模型,附基准分数和完整轨迹数据,适合评估开源方案与闭源模型的成本差距。

  2. elsewhere articles24

    心资本韩彦谈AI投资:泡沫之外,早期布局与非共识判断才是长期价值

    心资本创始合伙人韩彦在SuperReturn Asia 2026 AI & Deep Tech Investing Summit上表示,AI市场可能存在估值过热和泡沫,但AI仍是这个时代最具实质意义的技术变革之一。他以沐曦MetaX、曦望Sunrise等早期投资为例,强调从Day 0开始理解技术演进、坚持非共识判断,并指出未来只有既拥有长期数据积累又能用好AI的"1%"VC才能持续胜出。

  3. Thomas Wolf36

    “现在,获取关于 AI 公司内部情况的经过验证的信息,似乎尤为紧迫。”——@RyanGreenblatt

    引用Ryan Greenblatt@RyanGreenblatt

    I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.

  4. Deedy44

    Deedy Das 提出 Neolab 的多头逻辑:算力是 Helmer 式"垄断资源",当前存在获取算力与资本的窗口期,未来资金或算力价格可能恶化,竞争将更难。大实验室受创新者困境制约,难偏离编程现金牛、难自我蚕食收入、难追小于 $1B 的机会。许多 Neolab 已在产生可观收入,且人才认为其财务上行空间更大,至少 5-6 家大型潜在收购方。

    引用Deedy@deedydas

    The economics of a Neolab. A neolab is loosely defined as a startup of AI researchers who raises a lot of money pre-production to be able to finance GPU compute to take on a large AI problem. To buy 1000 GB300s or ~14 NVL72 racks will set you back $125-150M for 3yrs with 15-30% upfront. That’s about ~2-2.5MW. Thats about enough to do 10^25 flops a quarter and get to a GPT-4 level model which is 1-2 OOMs off frontier for pretraining. If you post-train on a great open source model, you have a better chance of getting to frontier. The risks are a) you need to spend millions on RL environments too and b) being lapped by another model release while being tied to a base model. For this to payback, you need to give your customers a better and ideally cheaper inference service than a base model and serve them for long enough to recoup your large investment. Even at 50% margin on inference, to recoup $10M in training means serving ~10T tokens (!) if you price like Fable / Astra given a standard cache read / input / output split ($2/M blended). And you have to justify being better than a release like Opus 5.5 which is even cheaper. Often, you end up charging your customers a huge premium in terms of platform fees and compute fees on top of pure inference. Meanwhile, every hour you’re not utilizing your GPUs you are burning money so you typically resell this compute back to a broker or run inference for open models / resell spot instances. At below a ~60% utilization on spot, you will still lose money. Add to that insane cost of talent. So what can you do with the compute? - Not play the model game at all. - Play an entirely different model game (Jev, World Labs) that if big labs played, would either a) cannibalize their business or b) be incrementally not significant revenue c) would cause too much distraction from the main main thing - Acquire a proprietary data set (Peridodic Labs) in enough volume in a domain of usefulness to eclipse frontier quality. Often happens in robotics, biology, chemistry. If you do overcome the challenge of building a model that is useful and well priced beyond big labs models, given the huge price of compute, you still need to play in an area where the revenue / compute ratio is signficant and market demand is large enough to payback your compute spend. It is a difficult game.

  5. elsewhere articles41

    群核科技用空间智能重建永泰龟城,24 亿高斯点刷新全球 3D 高斯重建纪录

    群核科技联合博主「特能斯」4 人 4 天在甘肃永泰龟城采集 6 万多张照片,通过 3D 高斯重建平台 Aholo Reality 生成 24 亿高斯点、60 万平方米的数字古城,刷新全球公开可查的 3D 高斯重建纪录。该数字场景已交给景泰当地文旅作为永久数字文化档案保存,并可通过 Aholo Reality 平台在线漫游;同一组 3D 场景还被用于 LuxReal 生成 AI 短剧《永泰无战事》。

  6. Hugging Face Daily Papers37

    Endless Exam:面向超级智能的数学构造基准,覆盖 14 类参数化问题族

    研究者推出 Endless Exam 基准,覆盖 14 个参数化数学构造问题族,用可验证的相对质量分数衡量模型在已发表数学前沿之前与之后的进展,且不将改进上限封顶为 1。在 69 个实例上评测 9 个模型,连续质量分数能区分表现差异,但 30 个已发表前沿参考无一被超越。该基准已开源生成器、验证器、参考基线、模型回答与分析,并新增 Claude Opus 5.5 评测。

  7. OpenRouter Announcements61

    OpenRouter 发布 Security Center 集中管理并降低 API key 风险

    OpenRouter 推出 Security Center,可在设置 > Security 下跨工作区查看所有 API key、识别风险 key,并支持一次最多 500 个批量禁用、归档或设置消费上限,所有套餐可用。

    推荐理由:OpenRouter 以自家 85 名员工 1000 多个 key 的审计为参照,介绍了安全中心的清理建议、风险评分和批量操作方法。

  8. Hugging Face Daily Papers42

    Tex-Zero:无需 3D 资产也能训练原生 3D 纹理生成模型

    研究团队提出 Tex-Zero,证明高保真原生 3D 纹理生成框架无需 3D 资产即可训练。该方法将高质量 2D 图像表示为 3D 空间中的平面,并通过分块随机旋转与聚合构造复杂几何结构,生成训练样本。基于这些数据训练的 Tex-Zero VAE 和 Tex-Zero DiT 在未见过真实 3D 资产的情况下,仍能生成细节精细的高保真 3D 纹理。

  9. Hugging Face Daily Papers35

    FAR:面向世界模型的自适应多线索情景记忆召回框架

    研究者提出 Future-Aware Recall(FAR),一个从未来感知预测监督中学习情景记忆召回的框架,通过负扩散预测损失近似条件对数似然来衡量记忆的预测效用,并训练一个推理时对未来盲的检索器。该检索器能学习各线索的相关性,自动决定每个查询该信任时间、姿态、视觉、音频中的哪些线索。在三种互补设置下,FAR 即使使用相同检索线索也优于手工设计的召回方法。

  10. Simon Willison54

    Simon Willison 发布 Bluesky 回复机器人检测工具

    Simon Willison 用 Opus 5.5 编写了一个 Bluesky 回复机器人检测工具,通过分析打字速度、发帖时间、互动模式等行为信号判断账号是否为自动回复机器人。工具展示每项测量和规则及触发信号最多的示例回复;作者称 Bluesky 开放 API 使此类调查比 Twitter 更可行,检测信号包括秒级连发回复、从不发布原创内容、专门回复高粉丝用户及使用问号等。