跳到正文
原文
Greg Brockman· @gdb · X·· 16 天前AI 评分64
AI 导读

Databricks 将 Astra 推广到全部约 3500 名工程师。其内部试点约 200 人的数据显示,Astra 在复杂任务上明显优于此前最强的 Opus 5 和 Sol 5.6,使用 Astra 的工程师整体编码支出增加约 60%,但在中低复杂度任务上相比既有模型提升不明显。团队通过 Unity Gateway 做分群实验,并为 Astra 单独设预算,鼓励工程师在复杂任务上选用 Astra、日常任务用更低成本模型;因数据保留政策尚未广泛部署 Fable,暂无 Astra 与 Fable 的可靠对比。

正文

wall-to-wall deployment of astra for engineers at databricks:

引用Patrick Wendell@pwendell
Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
在 X 查看被引用的帖子

来源:Greg Brockman · x.com