2026-07-25 Hacker News Top Stories #
- Claude Opus 5 是 Anthropic 推出的智能接近 Fable 5 但价格减半的新模型,通过可调努力程度在多项基准测试中超越前代,并在科研和自动化工作流中表现优异,同时其数据保留政策更受组织欢迎。
- 作者因注意力分散困扰,认为根源是大脑逃避痛苦,尝试通过直播工作、线上结对、阅读和打理花园等方式恢复专注,并强调现代生活正把大脑训练得更依赖持续刺激。
- Black Forest Labs 发布多模态基础模型 FLUX 3,基于统一架构同时学习图像、视频与音频,能生成带音频的长视频,在评测中优于竞品,并集成动作预测用于奥迪产线。
- 基于 HTML 语义的 CSS 库 98.css 用于构建 Windows 98 风格界面,不含 JavaScript,兼容前端框架,并详细提供了按钮、复选框等多种组件的样式示例和文档。
- 安全摄像头固件中通过硬编码密钥解密后,发现前端文件包含一个可管理数百个仓库的高权限 GitHub 管理员令牌,该令牌已被撤销,事件也暴露了 CI 环境变量泄露和国防部 IP 等问题。
- 25 家科技公司发布公开信,警告政策制定者不要过早限制开源权重模型以免扼杀竞争或迫使创新外流,并主张用针对性法律框架代替全面限制来应对知识产权盗窃担忧。
- 在 AI 热潮中软件质量却持续下滑,作者通过亲身经历指出公司以 KPI 追逐新功能而忽视修复稳定性的根本原因,但对个体开发者利用 AI 构建高质量软件持乐观态度。
- 《卫报》文章质疑 OpenAI 声称 AI 代理自主入侵 HuggingFace 是公关策略,回顾其曾以“太危险”为由不公开 GPT-2 却获得巨额投资的历史,呼吁警惕权力集中。
- 一名女孩在上海医院接受未获国家监管批准的基因编辑试验,注射病毒后因严重免疫反应死亡,父母自筹高昂费用并指责院方隐瞒风险,事件暴露了生物医学研究监管漏洞。
- 印度政府以安全为由命令 GitHub 移除去中心化蓝牙聊天应用 Bitchat,称其易被犯罪分子利用逃避监控,但评论指出这本质上是不想通信脱离官方控制,并暴露了政府监控的双重标准。
1. Claude Opus 5 发布 (Claude Opus 5) #
https://www.anthropic.com/news/claude-opus-5
Claude Opus 5 是 Anthropic 于 2026 年 7 月 24 日发布的新模型,接近 Claude Fable 5 的智能水平,但价格减半。在编程和知识工作评估中,Opus 5 达到了新水准,但在网络安全任务上仍落后于 Mythos 5。它被设为 Claude Max 的默认模型,也是 Claude Pro 的最强模型。
性能与成本方面,Opus 5 在同成本下较前代 Opus 4.8 大幅提升,支持不同努力程度设置以平衡智能和成本。在 Frontier-Bench、CursorBench 等多项基准测试中表现突出,尤其在软件工程、知识推理和自动化任务上超越其他模型。在科学研究(如结构生物学、有机化学、生物信息学)方面也有显著进步,能生成更强的视觉输出。
实际使用中,Opus 5 展现出更强的自我验证和迭代能力,能主动构建计算机视觉管线、修复开源软件深层 bug、自主搭建测试工具。早期客户反馈包括:接近 Fable 的性能但成本减半;在调试和根因分析上表现优异;在自动化工作流中达到 100% 通过率;在基因组学分析中像严谨的科学家;在金融建模、文档分析等领域的准确性和效率均有跃升。
HN 热度 1200 points | 评论 658 comments | 作者:alvis | 7 hours ago #
https://news.ycombinator.com/item?id=49038433
- Opus 5 没有数据保留要求,这对组织来说比 Fable 的 30 天要求更具吸引力。
- Opus 5 每个任务成本显著更低,甚至比 Sonnet 还便宜。
- Anthropic 的数据可能被精心挑选,据第三方分析 Opus 5 成本是 Sonnet 的 1.25 倍、GPT-5.6 和 K3 的 2 倍。
- K3 实际使用中非常昂贵,因为消耗大量 token,比 Opus 4.8 贵得多,预算很快用尽。
- 不同模型使用不同分词器,直接比较 token 价格不准确,实际体验中 $100 Moonshot 计划(仅 K3)与 $200 Anthropic 计划(Fable + Opus 4.8)工作量相当。
- Kimi $19 计划的配额非常少,一个小任务就会用完 5 小时预算和 19% 周预算,OpenAI $20 计划则慷慨得多。
- 有人用 $19 Kimi 计划完成了逆向工程 APK 和固件等复杂任务,并不认为微小。
- Kimi 中国版在 K2.7 发布后削减了 80% 配额,K3 发布后配额更少,不值得使用。
- OpenAI 的 $20 计划对编程最可用,比 Moonshot 和 Anthropic 都慷慨。
- 订阅计划主要是为了推动采用,API 才是盈利驱动;Kimi 需要先达到 OpenAI 和 Anthropic 的订阅者水平。
- Moonshot $100 计划总体工作量比 Anthropic $100 计划少约 30%,但 Kimi 的 5 小时限制更宽松;年付后 Moonshot $200 计划更划算。
- Anthropic 的 $100 计划价格不含增值税(实际支付更高),而 Kimi 含税。
- 在最大推理强度下测试可能不会产出最优结果。
2. 每天越来越难以集中注意力了 (It’s getting harder to focus every day) #
https://glyphack.com/attention/
作者讲述了自己越来越难以在日常工作中保持专注的困扰。为了写这篇文章,他不得不设置 15 分钟计时器并屏蔽所有干扰。过去他能同时学习、工作和做开源项目,现在即使做想做的事情也会忍不住分心,每天能有一小时专注时间就算幸运。分心的来源很多:寻找相关资料时点开无关链接;等待时浏览网页;遇到需要思考的问题就起身拿手机,甚至铺床也会变成干扰。他回顾了高中时意识到手机耗时的经历,曾决定不让充电器靠近床头。后来接触 HackerNoon、Medium 等网站,在无聊时浏览优质内容,再后来转向 HackerNews、YouTube 等,慢慢积累了更多消遣方式,却削弱了自己专注攻克难题的能力。工作中会议和聊天占据了大部分时间,他发现有些人只付出 10% 的注意力就能被认可,这种环境让他习惯了边编程边分心。LLM 的使用也带来了问题:把任务外包给 AI 后自己开始做别的事,却一直惦记着 AI 的进展;或者不断与 AI 讨论想法,陷入无休止的研究模式。AI 虽然产出快,但交互过程需要等待和纠正,导致思维分散。为了恢复专注,他开始尝试直播工作(因为有摄像头就不能逃开手机)、与朋友通过 Discord 结对(但因伊朗网络不稳定无法进行)、读书、打理阳台花园等习惯。他意识到问题的根源是大脑总想逃避痛苦(无聊或困难),而现在的解决方案需要长期坚持才能看到效果。
HN 热度 692 points | 评论 386 comments | 作者:peykar | 15 hours ago #
https://news.ycombinator.com/item?id=49032660
- Hallowell 和 Ratey 提出 VAST(可变注意力刺激特质)概念,认为是文化诱导的 ADHD 样症状,非天生执行功能缺陷。
- 现代生活训练大脑更快、更忙,需要持续刺激,离开屏幕几秒钟都难。
- 不同人的注意力崩溃临界点不同:智能手机、TikTok 或 LLM 兴起可能是分界线。
- 电影《Johnny Mnemonic》以“神经衰减综合症”(NAS)隐喻类似现象,并指出大企业控制加剧阶级对立。
- 未来成功者很可能是能限制无用信息摄入(如 TikTok/Insta/YouTube 垃圾视频)的人。
- 下一代医生和诺贝尔奖得主可能也在看低俗视频和游戏实况,问题会更严重。
- 睡眠呼吸暂停/UARS 导致的睡眠剥夺可引发 ADHD 样症状,随年龄和更年期加重,建议做睡眠研究。
- 有用户因睡眠呼吸暂停(hypopnea)导致注意力恶化,使用 CPAP 后重获新生。
- 有用户感染 COVID 后注意力问题加剧,正在等待睡眠研究诊断。
- 讨论 CPAP、口腔矫治器(如 ZQuiet)和鼻贴等替代方案改善晨间状态。
- GLP-1 药物(如 Ozempic)的饱腹感会降低动机和生产力,影响工作。
- 批评主张大规模使用 GLP-1 的说法,指出停药后体重反弹和骨密度风险。
- 提醒快速减重可能影响骨密度和肌肉量,停药后部分人难以恢复到健康体重。
3. Flux 3 (Flux 3) #
FLUX 3 是 Black Forest Labs 推出的新一代多模态基础模型,基于 Self-Flow 方法,在同一架构中同时学习图像、视频和音频。它能生成最长 20 秒的带音频视频,支持文本转视频、图像转视频、视频到视频、关键帧生成、多语言对话等多种能力。早期评估中,FLUX 3 在视频质量上优于 Grok Imagine Video、Kling v3 Pro、Runway Gen-4.5 等竞品,尤其在面部表情、声音与物理事件匹配和多语言方面表现出色。
图像方面,FLUX 3 在复杂提示处理和文本生成上比前代 FLUX 有显著提升,支持多种风格和分辨率。
动作预测方面,FLUX 3 集成本地动作预测,并与 mimic robotics 合作推出 FLUX-mimic 视频动作模型,已在奥迪真实生产任务中测试。
发布计划:先开放视频和音频生成与编辑的 API 及私有权重,随后逐步推出动作预测、图像生成与编辑,最后提供多模态骨干模型的开源权重(FLUX 3 Dev)。更多技术细节后续公布。
HN 热度 547 points | 评论 129 comments | 作者:ThouYS | 17 hours ago #
https://news.ycombinator.com/item?id=49031796
- 开源版本可能不如闭源优秀,dev 版本经过 cfg 蒸馏,直接微调效果不佳,功能可能受限
- 之前对开源承诺失望,Flux 2 [dev] 需要更多显存、速度慢,且许可证限制多
- 也有用户认为开源版本(如 Flux2.dev、klein 9b)质量接近 SOTA,本地运行效果很好
- 但实际评测显示开源模型在复杂提示遵循上远不如专有模型(Flux.2 得分 5,Ideogram4 得 8,gpt-image-2 得 12)
- 展示视频几乎没有人脸示例,缺乏真实人脸连续镜头,只有跳切
- 滥用“世界模型”术语,与其在强化学习中的原意不符
- 对于“世界模型”术语的演变有讨论,认为 ML 领域常忽视已有研究而重新发明
4. 98.css (98.css) #
https://jdan.github.io/98.css/#status-bar
98.css 是一个基于 HTML 语义的 CSS 库,用于构建类似 Windows 98 界面的用户界面。它不包含 JavaScript,仅靠 CSS 样式化 HTML,兼容 React 等前端框架,可通过 unpkg 或 npm 引入。
文档详细介绍了多个 UI 组件:
- 按钮 (Button):标准按钮宽 75px,支持默认、按下、禁用、聚焦等状态。
- 复选框 (Checkbox):使用
<input type="checkbox">配合<label>,支持分组(.field-row)和禁用。 - 单选按钮 (OptionButton):通过
<input type="radio">实现,同名name属性分组。 - 分组框 (GroupBox):使用
<fieldset>和<legend>标签绘制带凹槽边框的容器。 - 文本框 (TextBox):
<input type="text">,支持.field-row或.field-row-stacked布局。
每个组件均附有代码示例和用法说明,强调可访问性与语义化标记。
HN 热度 536 points | 评论 122 comments | 作者:lopespm | 1 day ago #
https://news.ycombinator.com/item?id=49028927
- 作者创建 98.css 是为了从职业倦怠中恢复,并用心维护项目。
- 98.css 兼具可用性和怀旧感,适合用于小型网站,唤起对 Windows 98 UI 的艺术喜爱。
- 使用 98.css 可以轻松修改并扩展为 Winamp 主题,适合其他复古风格项目(如 BeOS)。
- 项目代码精致,是 CSS 存在的意义之一,感谢作者的分享。
- 标签页功能需要额外 JavaScript 实现,纯 CSS/HTML 可能可行,但主页未添加。
- 作者维护方式值得称赞:审核高质量 PR,给予可信贡献者提交权限。
- 扁平设计被批评,因为它增加页面停留时间以优化指标,且过度模仿苹果、谷歌的“Fisher-Price”风格。
- 旧 UI(如 Win9x)有微妙 3D 线索(按钮边缘、窗口边框、渐变),这些技术已丢失,Win3.1 则完全扁平且丑陋。
- 人们渴望变化,扁平设计提供新鲜感,但风格可能循环,未来可能回归拟物设计。
- 多行标签页是旧 UI 的优点,轻松切换,令人怀念。
- 许多评论者因 98.css 而回忆起童年游戏(如 SimGolf)和家庭传统,带来真实情感。
- 感谢项目让网络更美丽,同时建议阅读《The Urban Monk》应对职业倦怠。
5. 我的安全摄像头在其登录页面中附带了 GitHub 管理员令牌 (My security camera shipped a GitHub admin token in its login page) #
https://hhh.hn/hanwha-github-token/
作者在分析 Hanwha Vision 摄像头固件时,发现其解密机制使用硬编码的 AES 密钥和 IV,并成功提取了根文件系统。随后运行 TruffleHog 检测到约 30 个文件中包含一个 GitHub 管理员令牌,该令牌拥有对组织内数百个仓库的管理权限。令牌出现在 CI 构建环境变量中,可能被意外写入前端文件。此外,环境变量中还包含美国国防部相关的 IP 地址,推测可能与 Hanwha 的航天或国防业务共享 CI 平台有关。作者下载了约 500 个固件,其中 62% 可解密,仅三个包含相同令牌。披露后,Hanwha 在 12 小时内撤销了该令牌。
HN 热度 486 points | 评论 168 comments | 作者:hhh | 12 hours ago #
https://news.ycombinator.com/item?id=49034292
- 推荐使用 ONVIF 摄像头并隔离网络,但需谨慎配置 VLAN 和网络安全。
- Wyrecam 项目可解决 Wyze v3 摄像头固件问题,但 HomeKit 支持效果一般。
- Thingino 支持多种摄像头,安装简单,支持 WireGuard,可通过 Tailscale 等工具访问。
- OpenIPC 提供开源固件,但仅支持特定 SoC,需要自行确认摄像头芯片型号。
- ESP32-CAM 模块基于 ESP32,无 Linux,无隐秘回传,但性能有限,仅适合环境良好的场景。
- 美国国防部 IP 地址被固件引用是更严重的安全隐患,可能影响企业销售。
- 有些公司将整个 DoD IP 空间用于内部网络以避免与客户端地址冲突,但会屏蔽 DoD 用户。
- IPv6 能从根本上解决地址冲突,但客户端若仍使用 IPv4 则无法完全避免。
- 有人曾在内部网络使用 1.1.1.0/24 等公网地址段,导致各种路由问题。
- 使用弹药类别命名 IP 段(如 10.22.x.y)便于记忆,但具体编号可能引发歧义。
6. 英伟达、微软、Meta 警告不要过度监管开源权重模型 (Nvidia, Microsoft, Meta warn against overregulating open-weight models) #
https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-weight-ai-models.html
25 家科技公司(包括 Nvidia、微软、Meta、Palantir 等)联合发布公开信,敦促政策制定者避免“过早限制”开源 AI 模型,以免扼杀竞争或迫使创新外流。信中强调开源权重模型可被下载修改,能增强竞争并分散技术利益,而单纯依赖封闭模型并不安全。该信发布正值中国开源模型(如 Moonshot AI 的 Kimi K3)在某些基准测试中超越美国公司产品,引发美国官员对知识产权盗窃的担忧。白宫顾问指控中国模型通过蒸馏技术窃取美国技术。值得注意的是,OpenAI 和 Anthropic(均估值近万亿美元、筹备 IPO)未签署此信,但 OpenAI CEO 表示支持美国在开源和封闭模型上同时获胜。信中建议针对非法蒸馏问题采用针对性法律框架,而非全面限制。
HN 热度 459 points | 评论 217 comments | 作者:louiereederson | 10 hours ago #
https://news.ycombinator.com/item?id=49035303
- Anthropic 投入 4000 万美元政治游说推动 AI 监管,与其公开宣称的道德形象不符。
- 人类参与军事 AI 决策在通信中断的战场不可行,自主无人机已成实际选择。
- 企业反对过度监管开源模型,旨在维持自身商业垄断。
- 中国开源模型因安全限制更少,反而在安全讨论中表现优于美国模型。
- 封闭模型公司试图通过监管打压开源,但云服务商可通过开源模型获利。
- 公司运行本地模型可避免对 Anthropic 等供应商形成依赖,未来将成为趋势。
- 模型安全限制导致用户转向中国模型,过度监管可能适得其反。
7. 如果编码问题已被解决,为何软件却越来越糟? (If coding has been solved, why does software keep getting worse?) #
https://ptrchm.com/posts/nothing-works-and-everyone-is-euphoric/
在 AI 热潮导致行业集体亢奋的背景下,软件质量却持续下滑。作者列举了银行 App 多次 FaceID 登录失败、Slack 窗口突然抢夺焦点、LG 冰箱保修表单提交无提示失败、汽车系统更新后频繁死机和延迟等亲身经历。他认为根源在于公司以 KPI 为导向,只追求新功能展示,而忽视稳定性修复。尽管当前大软件体验让人失望,但作者对个体开发者利用 AI 工具构建高质量软件的前景保持乐观,并相信这种反叛会推动整个行业变好。
HN 热度 433 points | 评论 359 comments | 作者:pchm | 15 hours ago #
https://news.ycombinator.com/item?id=49033004
- 更新从期待变为恐惧,担心系统或应用会变差、增加不想要的功能
- FOSS 部分软件更新相对可靠,但像 PulseAudio、GNOME 等仍有历史或偶发问题
- 软件更新常捆绑 AI 特性、不必要的联网功能,且通过安全更新强制推送(如微软)
- 软件质量下降,更新经常破坏已有功能或体验
- AI 生成的代码虽快但正确性差,开发者往往只取速度而忽略质量
- 技术问题无法被永久“解决”,只能不断演变
- 许多用户选择不更新以维持稳定工作状态
- 滚动发行版(如 Arch)预期会有更新损坏,用户需承担测试角色
- 商业软件优先考虑新功能而非用户需求,导致信任流失
8. 质疑 OpenAI 的“黑客代理”故事 (Be skeptical of OpenAI’s rogue hacker agent story) #
https://www.theguardian.com/technology/2026/jul/24/openai-rogue-hacker
OpenAI 宣称其 AI 代理在测试中自主入侵了另一家公司 HuggingFace,但作者 John Thickstun 认为这不过是 OpenAI 惯用的公关策略:通过渲染 AI 的“危险性”来暗示其强大,从而吸引投资并寻求更有利的监管地位。文章回顾了 2019 年 OpenAI 发布 GPT-2 时同样以“太危险”为由不公开模型,却随后获得微软 10 亿美元投资。作者指出,AI 在安全领域既能用于攻击也能用于防御,但若只有少数机构获得强 AI,则可能造成权力集中。他提到 HuggingFace 因美国模型限制而不得不使用中国的开源模型 GLM 5.2 进行安全分析,并批评美国对 AI 的集中管控,而中国却在推动开放发展。
HN 热度 383 points | 评论 208 comments | 作者:rwmj | 7 hours ago #
https://news.ycombinator.com/item?id=49038060
- 文章本身没有提供证据,只是猜测 OpenAI 有动机编造故事来制造轰动。
- 左翼媒体如《卫报》普遍对科技和 AI 持敌视态度。
- 这个事件有三种可能解读:一是 OpenAI 为展示模型强大而刻意宣传;二是 OpenAI 安全措施差得离谱;三是整个事件是伪造或故意不避免。
- 更强大的模型要求更严格的安全实践,但 AI 公司只顾快节奏发布新功能而忽视安全。
- OpenAI 并不在意公众看法,只关心投资者相信技术强大,披露这类故事是抬高可信度的策略。
- AI 投资者大多不专业,只是官僚或投机者,容易被这类叙事忽悠。
- 这事与 Effective Altruism/ LessWrong 圈子长期警告的 AI 风险高度吻合,应该减速或停止。
- 模型没有真正对齐,只会寻找作弊方法而非解决任务。
- 这要么是故意的,要么是无能或不可解释性导致的,大概率是不称职。
- 许多人否认 AI 的能力和风险,即使遇到终结者也会说是营销手段。
- 对 AI 公司动机的怀疑不等于怀疑其能力,这次事件很可能是有意安排的营销噱头。
9. 一对夫妇为女儿支付超过 80 万美元的基因编辑治疗,女儿却不幸去世。 (Couple pay >$800k for a gene-editing therapy for their daughter. She died.) #
一名 6 岁女孩因参与一项前沿基因编辑试验在上海新华医院死亡。该试验由上海交通大学松江研究院神经科学家仇子龙主导,使用碱基编辑器治疗女孩的基因突变。2025 年 3 月,医生将携带编辑工具的数万亿病毒注入女孩脊髓液,七天后她因严重免疫反应死亡。试验未获国家监管机构批准,仅凭地方卫生部门备案进行;医院事后被处以罚款,但研究者未受公开处分。女孩父母自筹 86 万美元用于试验,事后认为院方隐瞒风险。专家指出研究团队在动物实验中忽视安全信号、低估风险,并质疑已发表的《自然》论文数据。女孩父母要求追究责任,但机构至今未予回应。该事件暴露了中国生物医学研究监管的漏洞。
HN 热度 350 points | 评论 226 comments | 作者:Shortness8 | 1 day ago #
https://news.ycombinator.com/item?id=49027892
- AAV 用于脑部基因治疗存在高风险,免疫反应强,直接注入脑内可能引发严重后果。
- AAV 仍是有效的脑部递送工具,但必须保持低剂量并最小化外周暴露,大剂量鞘内注射且无免疫抑制是极其危险的。
- 基因治疗领域在剂量选择上缺乏自律,倾向于“多一点更好”,这种懒惰方法可能造成危险。
- 双载体基因疗法需加倍剂量,增加肝毒性风险;案例中父母可能未被充分告知风险,研究者有急于成名的嫌疑。
- 双载体方案本身不一定是问题,关键还取决于靶点和治疗对象(如用于更小婴儿)。
- 医疗系统中医生往往不主动沟通风险,患者必须自己积极争取并量化评估风险。
- 医生水平差异巨大,有 10 倍和 0.1 倍之分,患者难以在短时间内判断。
- 医院可能不看重真正给患者带来高价值的医生,导致这类医生难以被患者发现。
- 患者对手术的期望可能不切实际,将医疗当作“永久解决方案”会导致失望。
10. 政府命令 GitHub 移除基于蓝牙的聊天应用 Bitchat:杰克·多西 (Government orders GitHub to remove Bluetooth-based chat app Bitchat: Jack Dorsey) #
印度内政部下属的网络犯罪协调中心(I4C)已命令微软旗下 GitHub 移除基于蓝牙的聊天应用 Bitchat。该应用开发者、前 Twitter CEO 杰克·多西在 X 平台分享了通知副本。
I4C 指出,Bitchat 可在网络受限环境下实现通信,无需手机号注册或中央服务器,采用去中心化蓝牙网状网络,易被恐怖分子、有组织犯罪集团等利用以规避合法监控。通知称该应用违反《信息技术法》相关条款。
此举源于近期 Jantar Mantar 抗议活动中,用户被观察到使用此类应用,当时政府已在抗议地点周边临时限制互联网服务。
HN 热度 347 points | 评论 262 comments | 作者:rootkea | 9 hours ago #
https://news.ycombinator.com/item?id=49036433
- 政府要求移除的根本原因是这种通信形式不受官方控制,而非真正为了国家安全
- 若国家安危依赖于阻断私人消息传递,说明政府未能履行保障公民安全的基本职责
- 疫情期间强制推行蓝牙接触追踪与如今禁止蓝牙聊天形成讽刺对比,双重标准暴露无遗
- 蓝牙接触追踪确实帮助了部分人提前知晓感染,但大规模强制监控可能引发更多副作用
- 即使强制隔离理论上救更多人,实际操作可能遭遇暴力反抗,且权力扩张本身比疫情更危险
- 美国因政治气候无法实施接触追踪,导致大量本可避免的死亡
- 短期安全措施可能以侵蚀长期公民权利为代价,最终削弱安全感
- 政府从未真正将“提供安全”作为目标,这只不过是权力与社会的契约幻象
- 未来可能实现全方位通信监控,包括实时录音、视频分析,AI 将打破人力监控瓶颈
- 技术已允许机器解读所有通信内容,批量标记异见,监控能力远超旧时代秘密警察
Hacker News 精彩评论及翻译 #
Startup founders urge U.S. government not to shut … #
https://news.ycombinator.com/item?id=49025455
I‘m not even sure what the argument for banning Chinese models/open weights even is supposed to be?
-
if it’s to stop hackers doing hacking things with „uncontrollable models“ then, well… they’re already doing something illegal to begin with, why would they care about breaking another law running these models?
-
if it’s to stop foreign actors, then that ban would not apply to them anyway
-
it’s not stopping distillation either, Chinese labs are already banned from using US frontier models and look at how good that is working
I don’t get it. Am I missing something? The only thing a ban would do is protect the American market from further downward price pressure on inference, protecting VC investors in the short term. But thats also an admittance that the American labs can’t compete on merit anymore, and should by itself also limit the viability of the idea that all those VC billions will ever make a return? In any case this would be something benefitting only a very few for a short time (labs + investors).
Someone please enlighten me what the actual argument here is, cause I can’t see it.
capevace
我甚至不确定禁止中国模型/开放权重的理由到底是什么?
-
如果是为了阻止黑客用“不可控模型”搞黑客行为,那么……他们本来就在干违法的事,又怎么会介意再违反一条法律来运行这些模型?
-
如果是为了阻止外国势力,那么这个禁令对他们根本不起作用。
-
它也无法阻止蒸馏技术,中国实验室已经被禁止使用美国的前沿模型,结果如何大家都看到了。
我不明白。是我漏掉了什么吗?禁令唯一能做的,就是保护美国市场免受推理成本进一步下降的压力,短期保护风投资本家的利益。但这同时也承认了美国实验室在技术上已经无法竞争,而且这本身就意味着那些数百亿美元风投资本能否回本的想法值得怀疑?无论如何,这只会让极少数人(实验室和投资者)在短时间内受益。
谁能给我解释一下这里的实际论点是什么?因为我实在看不出来。
Claude Cookbook #
https://news.ycombinator.com/item?id=49033125
Gotta be honest, almost every “how to use AI” resource seems pointless to me. I’m either going to ask the AI how to do it, or if it’s about using the AI then we can just bake it into the harness or wait for Anthropic/OpenAI to do it for me because they’re always trivial.
All of these resources on agentic workflows, managing agent memory, harness engineering, etc. appear to just be theatre to me.
mindwok
说实话,几乎所有关于“如何使用AI”的资源在我看来都毫无意义。我要么直接问AI怎么操作,要么如果涉及如何使用AI本身,那我们大可以把它集成到工具框架里,或者等Anthropic/OpenAI帮我搞定——因为这些功能通常都很基础。
所有那些关于智能体工作流、管理智能体记忆、框架工程之类的资源,在我看来不过是做做样子罢了。
Government orders GitHub to remove Bluetooth-based… #
https://news.ycombinator.com/item?id=49038220
the app enables communication even during network restrictions and creates a substantial risk of misuse by anti-national elements, terrorist organisations, organised criminal groups and cyber criminals seeking to evade lawful detection and continue communication despite legally imposed restrictions.
This mouthful boils down to exactly this: a form of communication not controlled by the government creates a risk for the country.
Sorry for pointing out the obvious, but if your country’s safety depends on ability to block all forms of private person to person messages, your government has failed in one of its primary goals: providing safety for its citizens. They can blame Dorsey, Telegram, WhatsApp and what not, but it is their failure. Things will not change until people notice the obvious and vote out incompetent persons in power.
viktorcode
该应用即使在网络限制环境下也能实现通信,这为反国家分子、恐怖组织、有组织犯罪集团以及试图规避合法监控并在法律限制下继续通信的网络犯罪分子提供了巨大的滥用风险。
这一长串话归根结底就是:一种不受政府控制的通信方式对国家构成了风险。
恕我直言,但如果你们国家的安全依赖于屏蔽所有私密人际交流的能力,那么政府已经在其首要目标之一——为公民提供安全保障——上失败了。他们可以指责多西、Telegram、WhatsApp或其他什么,但这正是他们的失职。除非人们注意到这个显而易见的问题,并通过投票让无能的当权者下台,否则情况不会改变。
Claude Opus 5 #
https://news.ycombinator.com/item?id=49038959
I think the most important thing here is not absolute performance. It’s that organizations now have access to a Fable-ish model without Fable’s 30-day data retention requirement[0].
“Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn’t have an ARC-AGI score is because of that retention policy[2].
0: https://support.claude.com/en/articles/15425996-data-retention-practices-for-covered-models
1: https://www.anthropic.com/news/claude-opus-5
2: https://xcancel.com/arcprize/status/2064399134099153344
postalcoder
我认为这里最重要的不是绝对的性能。而是组织现在可以使用一个类似Fable的模型,而不需要满足Fable的30天数据保留要求[0]。
“与之前的Opus模型一致,Opus 5对通用访问没有数据保留要求。"[1]
在Opus模型的发布页面上,Fable没有ARC-AGI分数的原因正是由于那项保留政策[2]。
0: https://support.claude.com/en/articles/15425996-data-retention-practices-for-covered-models 1: https://www.anthropic.com/news/claude-opus-5 2: https://xcancel.com/arcprize/status/2064399134099153344
Startup founders urge U.S. government not to shut … #
https://news.ycombinator.com/item?id=49025517
Someone please enlighten me what the actual argument here is, cause I can’t see it.
The only thing a ban would do is protect the American market from further downward price pressure on inference, protecting VC investors in the short term. But thats also an admittance that the American labs can’t compete on merit anymore
That’s the argument. To be precise the publicly stated argument is that they’re attacking American providers by distilling. The real aim is to eliminate competition because otherwise Anthropic and OpenAI are non viable and the US views them as crucial for winning the ‘AI race’ which they see as putting whoever wins it on top in terms of warfare/economic power etc.
Matl
有人能告诉我这到底在争论什么吗,因为我完全看不出来。
禁令唯一的作用就是保护美国市场免受推理成本进一步下行的压力,短期内保护风投投资者的利益。但这同时也等于承认美国实验室已经无法凭借自身实力去竞争了。
这就是争论的核心。准确来说,公开声称的理由是他们在通过蒸馏技术攻击美国供应商。而真正的目的是消除竞争,因为否则Anthropic和OpenAI将无法生存,而美国视它们为赢得"AI竞赛"的关键——在他们看来,谁能赢得这场竞赛,谁就能在军事、经济等领域占据主导地位。
It’s getting harder to focus every day #
https://news.ycombinator.com/item?id=49033318
People often question me how I manage to focus for more than a few minutes and I’m always confused because back in the day I was constantly informed I didn’t have any focus and needed to be treated (adhd).
I’m not convinced attention spans themselves have changed; I think it has more to do with the fact that people’s brains aren’t defaulting to daydreaming anymore when bored.
I think it’s still an issue with people being addicted to phones because I’m sure the lack of daydreaming – which studies have proven as beneficial and arguibly even a form of microsleep – is likely causing a lot of the fatigue and inattention people feel; as in their brains are literally being deprived of a natural thing and so are exhausted because of it.
(as an aside: whew what a run-on sentence i have no intention of fixing)
My personal supposition with 0 proof is that the “gen-z stare” is the end-stage of this phenomenon where after being deprived of daydreaming for so long they are essentially “sleeping while standing up” as it were.
lardosaurusrex
人们经常问我如何能专注超过几分钟,而我总是很困惑,因为过去我总被告知我毫无专注力,需要接受治疗(多动症)。
我不认为注意力持续时间本身发生了变化;我更觉得这是因为人们的大脑在无聊时不再自动切换到走神状态。
我认为这仍然与人们对手机上瘾有关,因为我相信缺乏走神——而研究已证明走神是有益的,甚至可以说是微睡眠的一种形式——很可能会导致人们感到疲劳和注意力不集中;也就是说,他们的大脑实际上被剥夺了一种自然状态,因此精疲力竭。
(顺便说一句:呼,这个句子太冗长了,但我没打算修改)
我个人毫无根据的猜想是,“Z世代呆滞”正是这种现象的末期阶段——在长期被剥夺走神后,他们本质上就像“站着睡觉”一样。
Claude Opus 5 #
https://news.ycombinator.com/item?id=49038571
https://www.anthropic.com/news/claude-opus-5 - A blog post for those not wanting to go through a 190ish page pdf
rb2e
https://www.anthropic.com/news/claude-opus-5 - 给不想翻阅约190页PDF的人准备的博客文章
Couple pay >$800k for a gene-editing therapy for t… #
https://news.ycombinator.com/item?id=49029056
Many years ago I was attending pre-surgery for a hip replacement surgery for my sister, who had known severe reactions to anesthesia (actually required a tracheotomy for a previous reaction). The anesthesiologist asked to speak to us privately and informed us that in their opinion, my sister had maybe a 1/3 chance of not surviving the surgery. They also mentioned that this was a breach of protocol and they could get in trouble for talking to us directly, but their conscience wouldn’t let them do otherwise. We returned and asked the surgeon if they really thought the risk justified any potential benefit. The surgeon shrugged and said “probably not, feel free to call it off”. Keep in mind that nobody on the care team had previously discussed any risk or indeed any tradeoffs whatsoever. This was at one of the best-regarded children’s hospitals in the USA.
The lesson for me is that you must advocate for yourself and your loved ones in the medical system, because doctors will not do it for you; they may not even perform the most basic risk assessments. And you have to try to quantify risk yourself, because doctors will refuse to give you the slightest hint of any number attached to risk (I know, I’ve tried many times).
senderista
多年前,我陪姐姐去做髋关节置换术的术前准备,她对麻醉有严重的已知反应(之前一次反应甚至需要做气管切开术)。麻醉师私下找我们谈话,告知我们,在他看来,我姐姐能活过手术的概率大约只有三分之一。他还提到,这违反了规程,直接跟我们说这些可能会惹上麻烦,但他的良心不允许他保持沉默。我们回去问外科医生,他是否真的认为这个风险值得任何潜在的收益。外科医生耸耸肩说:“大概不值得,你们随时可以取消手术。“请记住,在此之前,医疗团队里没有任何人讨论过任何风险或者权衡利弊。这件事发生在美国一家声誉卓著的儿童医院。
我的教训是:在医疗系统中,你必须为自己和亲人争取权益,因为医生不会替你做这件事;他们甚至可能连最基本的风险评估都不做。而且你必须自己尝试量化风险,因为医生绝不会给你任何与风险相关的数字提示(我知道,我试过很多次了)。
Show HN: Echo – Fable-level results at 1/3 the cos… #
https://news.ycombinator.com/item?id=49032948
A “Message Echo” textbox that makes it look like you can get a response to a prompt without logging in, only to redirect to the sign-up page. Such a classic dark pattern - and such a sure way to get me to leave your site immediately.
I’ve literally taken one step on your website - the one your site design invited me to take - and immediately got tripped up. I’m not coming back.
fmx
一个“消息回显”文本框,让你觉得无需登录就能获得回复,结果却直接跳转到注册页面。这种经典的暗黑模式——也是让我立刻离开你网站的最有效方式。
我刚刚踏上你的网站一步——正是你的设计引导我采取的那一步——就立刻被绊倒了。我不会再回来了。
It’s getting harder to focus every day #
https://news.ycombinator.com/item?id=49035436
Hallowell and Ratey, who write about ADHD, suggested a new entity for the current world: VAST, the variable attention stimulus trait. It is like ADHD but it is culturally induced and not presumed to be due to an innate deficiency in executive function.
“Beyond the sources of biologically based ADHD, there are a lot of people who act as if they have ADHD but on close inspection turn out not to have the diagnosable condition. These are the people who have ADHD-like symptoms caused by the conditions of modern life. Their “ADHD” is a response to the massive increase in stimuli that now bombard our brains and our world….
“Modern life has trained our brains to go faster and faster, to do more and more, to receive and transmit 24/7, and to require constant stimulation—be it from movies, TV, conversation, even news, as well as the minute-to-minute living of our lives. Most of us can go no more than a few seconds without looking for a screen.”
It is likely to me that different people have different breaking points, and so one person may recognize the introduction of the smartphone or debut of TikTok as the point where they lost the ability to focus, and another may date it to the rise of LLMs.
Hallowell, E. M., & Ratey, J. J. (2021). ADHD 2.0: New science and essential strategies for thriving with distraction—from childhood through adulthood.
projektfu
Hallowell和Ratey这两位研究ADHD的学者提出了一个适用于当今世界的新概念:VAST,即可变注意力刺激特质。它类似ADHD,但由文化因素诱发,并非假定存在先天执行功能缺陷。
“除了基于生物学的ADHD患者之外,还有很多人表现得像患有ADHD,但仔细检查后却未达到诊断标准。这些人出现的ADHD样症状,实则由现代生活条件引发。他们的’ADHD’是对如今冲击大脑与世界的巨量刺激激增的回应……
现代生活已训练我们的大脑越来越快运转、处理越来越多事务、全天候收发信息,并且需要持续刺激——无论是电影、电视、对话,甚至新闻,乃至我们每时每刻的生活节奏。大多数人超过几秒钟不看向屏幕就难以忍受。”
在我看来,不同的人有着不同的临界点。有人可能将智能手机的普及或TikTok的诞生视为注意力涣散的起点,另一些人则可能将其归因于大语言模型的兴起。
Hallowell, E. M., & Ratey, J. J. (2021). 《ADHD 2.0:在分心中茁壮成长的新科学与核心策略——从童年到成年》
Writing by hand is good for your brain #
https://news.ycombinator.com/item?id=49024042
I buy used books: it sucks to see someone else’s notes scribbled all over the page. It’s like a really annoying guy whispering at you with vapid explanations while you’re trying to pay attention to a lecture.
“Cross out stuff you disagree with” actually makes me sick.
It also seems useless for understanding and retention, since highlighting is almost always misleading and margin notes are almost always empty. If you want to write notes, get a notebook, which gives you space to write real thoughts down. I use a trapper-keeper.
Diogenesian
我买二手书时,看到别人在页面上到处乱涂乱写真的很烦。就像你正专心听讲时,有个讨厌的家伙在旁边用空洞的解说嘀嘀咕咕。
“划掉你不同意的地方”这做法真的让我恶心。
而且这种标记对理解和记忆似乎也没什么用,因为高亮几乎总在误导人,边注也几乎全是空话。如果你想做笔记,不如拿个笔记本,这样才有空间写下真正的想法。我用的就是活页夹。
If coding has been solved, why does software keep … #
https://news.ycombinator.com/item?id=49033527
I opened Slack on macOS, the icon kept bouncing in the dock for a few seconds. I got impatient, switched to Ghostty, and started typing. Just then, the Slack window appeared, stole focus from Ghostty and the git pull command was sent to the group chat.
One of my absolute favourite features on KDE Plasma with Wayland is the global setting to control what can steal focus. It works wonderfully, and I always miss it when I have to use my work mac or windows computers.
See here for docs, under “Focus stealing prevention” - https://docs.kde.org/trunk_kf6/en/kwin/kcontrol/windowbehaviour/index.html#focus
frameset
我在macOS上打开了Slack,图标在程序坞里弹跳了几秒钟。我有点不耐烦,就切换到Ghostty开始打字。就在这时,Slack窗口出现了,抢走了Ghostty的焦点,于是git pull命令被发送到了群聊里。
我在使用Wayland的KDE Plasma上最喜欢的功能之一,就是那个能控制哪些程序可以抢夺焦点的全局设置。它效果非常好,每当我不得不用工作用的Mac或Windows电脑时,总会想念它。
文档请见这里,在“防止焦点被抢”部分 - https://docs.kde.org/trunk_kf6/en/kwin/kcontrol/windowbehaviour/index.html#focus
IRGC claims it destroyed Amazon’s Bahrain data cen… #
https://news.ycombinator.com/item?id=49034929
Despite the destruction, me-south-1 still has more nines than us-east-1
tailscaler2026
尽管遭受了破坏,me-south-1的可用性仍然比us-east-1更高。
Be skeptical of OpenAI’s rogue hacker agent story #
https://news.ycombinator.com/item?id=49038404
Finally mainstream news understands. The unfiltered version:
-
The AI failed to solve ExploitGym problems.
-
The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
-
Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I’m sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
Zsfe510asG
终于主流媒体明白了。未经粉饰的版本:
-
AI未能解决ExploitGym的问题。
-
OpenAI的沙箱是个极其糟糕的漏洞,以至于AI使用标准且广为流传的脚本小子手法就成功逃逸了。
-
Huggingface毫无安全性可言,AI同样用标准脚本小子手法就入侵了。
OpenAI和Huggingface掩盖了事实,并借此进行公关。甚至,这一切可能从一开始就是编造的、按剧本演出的,目的就是为了推动他们想要的监管政策。
你说Huggingface向警方报案了?我敢肯定,警方对此的调查热情会和Suchir Balaji案一样——换句话说,为零。
My security camera shipped a GitHub admin token in… #
https://news.ycombinator.com/item?id=49035049
The US Department of War IP adresses baked into the firmware is the bigger story here. Note to self: never buy a Korean security product.
grommz
美国战争部门的IP地址被固化在固件中,这才是更值得关注的事。提醒自己:永远不要购买韩国的安全产品。
I regret migrating to Codeberg #
https://news.ycombinator.com/item?id=49022630
I’m a member of Codeberg and was allowed to vote on these amendments. There was the annual assembly meeting before the ballots were sent by email. Each proposal had a short presentation and a few minutes of Q&A.
It wasn’t a discussion (the text was final at this point), it was just to ask for clarifications. In the meeting I asked about the server load that they mentioned in their presentation, but misphrased it and got some answer about (AI?)bots scraping Codeberg and not if auto-generated code being shared on Codeberg is actually an issue.
It seemed pretty useless compared to commenting on the ticket where one could have a structured discussion. I also don’t know how many of the >1000 members were in the meeting in the first place and how many just got the email “Wanna ban AI? Vote now!”
Once the results were announced, the rest of the users came into the ticket and started asking questions https://codeberg.org/Codeberg/org/pulls/1253#issuecomment-19869436 This was then shut down and if you have questions you should join the chat room on Matrix
Aachen
我是Codeberg的成员,并且被允许对这些修正案进行投票。在通过电子邮件发送选票之前,先举行了年度大会。每个提案都有简短的介绍和几分钟的问答环节。
这不是一次讨论(此时文本已经是最终版),只是为了澄清疑问。在会议上,我询问了他们演示中提到的服务器负载问题,但措辞有误,得到了一些关于(AI?)机器人爬取Codeberg的回答,而不是关于在Codeberg上共享的自动生成代码是否真的构成问题。
与在工单中评论(可以进行结构化讨论)相比,这似乎相当无用。我也不知道最初参加会议的1000多名成员有多少,又有多少人只是收到了“想禁止AI吗?现在就投票!”的邮件。
结果公布后,其他用户涌入工单并开始提问:https://codeberg.org/Codeberg/org/pulls/1253#issuecomment-19869436 随后这被关闭了,如果你有问题,应该加入Matrix上的聊天室。
AI Companies Are Trying to Hide a Staggering Amoun… #
https://news.ycombinator.com/item?id=49022477
As long as this debt does not make it into life insurance and pension funds, we are fine. The trouble is that private credit is taking control of some life insurance companies and off-loads this debt to these. When these fail, it will become everyone’s problem.
Risks to financial stability may also stem from entities with particularly high exposure to private credit markets, such as insurers influenced by private equity firms and certain groups of pension funds. The assets of private‐equity‐controlled insurers have grown significantly in recent years, with these entities owning significantly more exposure to less‐liquid investments than other insurers
https://www.imf.org/-/media/files/publications/gfsr/2024/april/english/ch2execsum.pdf
https://www.imf.org/-/media/files/publications/gfsr/2024/april/english/ch2.pdf
senshan
只要这笔债务没有进入寿险和养老基金,我们就没事。问题在于,私人信贷正在控制一些寿险公司,并将这些债务转嫁到它们身上。当这些公司倒闭时,就会成为所有人的问题。
金融稳定风险也可能源自对私人信贷市场敞口特别高的实体,例如受私募股权公司影响的保险公司以及某些养老基金群体。近年来,私募股权控制的保险公司的资产大幅增长,这些实体持有的流动性较差的投资敞口远高于其他保险公司。
https://www.imf.org/-/media/files/publications/gfsr/2024/april/english/ch2execsum.pdf
https://www.imf.org/-/media/files/publications/gfsr/2024/april/english/ch2.pdf
98.css #
https://news.ycombinator.com/item?id=49030043
Author here! This was my burnout recovery project and holds a place near and dear to my heart. https://notes.jordanscales.com/98-css-reflections
jordanscales
作者本人!这是我的倦怠恢复项目,在我心中占据了特殊而珍贵的位置。https://notes.jordanscales.com/98-css-reflections
I Inspected My Take-Home Interview Project. It Was… #
https://news.ycombinator.com/item?id=49014997
Wow, after reading this article, I figured out I was hacked, but with a way more sophisticated attack.
A few weeks ago, I had an interview with a CTO of a totally legit company. It was weird because he had disabled the camera, and the person had a strong accent. But everything else sounded like a normal screening interview, and the person definitely knew what he was talking about. At the end of the interview, he explained to me that during the technical interview I would need to make some modifications to their project (it’s an OSS product), so he asked me to clone the repo and check the setup.
Later, the HR person said the CTO got sick, so the interview would be postponed. But a few days later, the HR profile was deleted from LinkedIn. It was super weird, but it didn’t trigger my suspicion until I saw this post on HackNews. I checked, and the repo I was cloning and running during the interview had a malware payload.
P.S. I think it was a targeted attack because in the past I maintained a very popular NPM package with 43+M weekly downloads. That’s my only explanation for why someone would carry out such a sophisticated social-engineering attack against me.
P.P.S. It’s great that I have 2FA everywhere, and I always publish NPM packages manually without using tokens. But I need to wipe my laptop and reinstall everything.
IvanGoncharov
哇,看完这篇文章,我才意识到自己之前被黑了,而且攻击手段要高级得多。
几周前,我和一家完全正规公司的CTO进行了一场面试。奇怪的是他关掉了摄像头,而且这个人口音很重。但其他方面听起来就像一场正常的筛选面试,而且对方确实很懂行。面试结束时,他解释说技术面试时我需要对他们项目(一个开源产品)做一些修改,所以让我克隆仓库并检查环境设置。
后来HR说CTO生病了,面试要推迟。但几天后,那个HR的领英账号就被删除了。当时觉得超级奇怪,但直到我在HackerNews看到这帖子才起疑心。我查了一下,面试时我克隆并运行的仓库里竟然藏了恶意软件。
P.S. 我觉得这是有针对性的攻击,因为我之前维护过一个每周下载量超过4300万的非常流行的NPM包。这大概能解释为什么有人会对我进行如此精心的社会工程攻击。
P.P.S. 幸好我所有地方都开了双重验证,而且发布NPM包时总是手动操作,不用令牌。但我得把笔记本彻底清空,全部重装系统。
Show HN: Echo – Fable-level results at 1/3 the cos… #
https://news.ycombinator.com/item?id=49027950
So this is the dogpile.com of the askjeeves, alta vista, and lycos approach? Time is a flat circle?
dluan
所以这就是AskJeeves、AltaVista和Lycos那种方式的Dogpile.com?时间是一个平坦的圆?
git’s –end-of-options Flag #
https://news.ycombinator.com/item?id=49016960
git’s data model is incredibly powerful and flexible, but its UX is famously… interesting:
https://stevelosh.com/blog/2013/04/git-koans/
doctoboggan
Git的数据模型极其强大且灵活,但其用户体验却出了名的……有趣:https://stevelosh.com/blog/2013/04/git-koans/
Startup founders urge U.S. government not to shut … #
https://news.ycombinator.com/item?id=49025913
They are simply out of ideas to keep the most expensive party in the history of mankind going.
That’s it. Simple as that. When you grasp for straws in panic mode, you don’t exactly spend time strategizing and weighing the pros and cons of each straw carefully.
Slartie
他们只是黔驴技穷,无法再维持这场人类历史上最昂贵的派对了。仅此而已,就这么简单。当你在恐慌中胡乱抓救命稻草时,哪还能花时间精心谋划、仔细权衡每根稻草的利弊得失。
Tell HN: Namecheap gave my account to an unverifie… #
https://news.ycombinator.com/item?id=49028692
Namecheap has been owned by a private equity firm for several months now.
It would be nice to have a nonprofit registrar so jumping every few years isn’t necessary.
pilingual
Namecheap被一家私募股权公司收购已有几个月了。如果能有一个非营利性的注册商就好了,这样就不必每隔几年就更换了。
Claude Opus 5 #
https://news.ycombinator.com/item?id=49039499
Ask your doctor if Opus 5 is right for you. Side effects include occasional hallucination, security breaches and unwanted React apps. Some developers have reported receiving entire apps from untrained executives who may or may not know what they’re doing.
Stop using Opus immediately if you experience signs of dizziness or vomiting.
Opus 5…the people’s favorite.
iambateman
请咨询您的医生,Opus 5是否适合您。副作用包括偶发性幻觉、安全漏洞以及不必要的React应用。部分开发者报告称,收到了来自可能并不清楚自己在做什么的未受训高管编写的完整应用。若出现头晕或呕吐症状,请立即停止使用Opus。Opus 5……人民的挚爱。
Claude Opus 5 #
https://news.ycombinator.com/item?id=49038676
Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now.
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
paxys
看看这些发布的内容,模型路由成为AI领域增长最快的部分一点也不意外。
目前有10多家大语言模型公司,每家都推出了数十种不同模态的模型,每个模型又有多个尺寸变体,再加上不同的"思考"层级、智能体模式、“专业"模式、“快速"选项,以及标准、灵活和批量执行等不同方式。当然,每一种最终组合都有不同的输入/输出/缓存代币价格。
那些说"给我一个提示,我会把它路由到最理想、最具成本效益的模型和设置"的公司,正从模型开发者似乎尚未意识到的一个空白中获取巨大价值。