每日简报

2026-09-07

该源今日无内容。

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

👍 272

As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedbac

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

👍 164

Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how much it will matter later, a

Editable Visual Design

👍 42

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides pre

Iris: Climbing to the Search Frontier

👍 36

We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are reverse-constructed from the hyperlink structure of a web corpus: we author multi-hop chains over an entity graph distilled fro

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

👍 30

Agent performance depends jointly on the model parameters and the executable harness code that manages context and control flow. Optimizing either component in isolation can leave the system bottlenecked by its frozen counterpart: weight updates can change which harness is effective, while harness u

PACE: Towards Surfacing Hidden Conflicts in User Requests

👍 23

Personalized assistants should not only comply with user requests but also assess whether those requests are appropriate given the user's current circumstances. However, prior work has primarily focused on accurately executing requests, overlooking the need for assistants to account for context and

WorldReward: Reward Modeling for Camera-Conditioned World Models

👍 23

Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory executi

Using Grounded Theory for Agent Behavior Analysis at Scale

👍 18

Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, often unfamiliar tasks where pre-built classifiers fall short. We propose to bring grounded theory into agent trajectory analysis: a six-decade-old qualitative method from the soci

Environment Evolution for Terminal Agents

👍 18

Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments n

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

👍 16

An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a ca

Principia: Relational Physics Tests for Video Models

👍 16

Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the

该源今日无内容。

该源今日无内容。

佬们,电报自带的搜索怎么搜不到东西了又

佬们,前一阵电报自带的搜索还有#检索非常好用,基本上全部群聊(已加入的)的消息都能搜索到。这两天再用的时候怎么搜索结果很少了,寥寥无几。大家有什么好的替代方案吗? 2 个帖子 - 2 位参与者 阅读完整话题

bybit asia能否转eu

问一下论坛的各位佬,之前开了bybit的asia地区的卡,现在看到有佬友发了开通eu卡的教程,下卡好像还挺快,因为eu更优惠,想问一下可以转eu吗,还是说要注销账户之后重新注册呢? 2 个帖子 - 2 位参与者 阅读完整话题

闹鬼一样的宽带故障,装维解决不了.....

症状:从天刚蒙蒙亮开始,光信号就会逐渐下降,直至大约早上8点断网(每天都不一样,可能前后波动一两个小时),断网时信号降至-34db 来了两波师傅,第一次更换了入户光纤插头,第二次更换了分光器,故障依旧。 现在实在不知该如何跟运营商沟通了…感觉对方也尽力了,但问题就是没法解决 2 个帖子 - 2 位参与者 阅读完整话题

还是忍不住,把智谱的套餐升级到MAX

V2 Pro的5H有想法的时候很快会耗光,光等又打断思路,偶尔用用公司的token干还行,但也不是长久 最终还是升级到了V2 Max,平均每个月320多,5H不停跑也用不完 就是没啥想法的时候,就纯浪费 4 个帖子 - 3 位参与者 阅读完整话题

中转站代理到 claude\ Codex 和 codex-cli 的使用姿势对比

我的远程主机安装了codex 和claude , 想请教一下大家在使用的时候装了哪些插件。 比如,codex 的 cli 似乎会缺少很多桌面端的功能,比如 computer use , 或者一些网页搜索功能。 同一个链接喂给CLI, 它需要curl好多次才能拿到结果, 但是桌面端速度很快,显示的工具调用是“网页搜索”, 难道是需要装一些插件才能达到类似的效果吗? 在远程设备上做开发,codex-cli/codex/claude-cli 哪个更友好一些呢? 1 个帖子 - 1 位参与者 阅读完整话题

zed agent 今天能使用吗

今天zed agent一直用不了,今天问个问题一直打圈圈,没有任何回复。 升级到最新版,不行 太奇怪了,各位使用情况如何。 前几天 学生认证 过,确定可以使用,今天不知道啥原因,请教各位大佬。 1 个帖子 - 1 位参与者 阅读完整话题

fable5.1记忆问题

为什么感觉fable最近记忆有很大问题啊,老是乱记东西,比如前面和他订好了,什么时候用grok,什么时候用glm,结果刚刚让它调用gpt6来做一个类似设计skill的工作,后面发现它只调gpt了,一问他发现他把xxx要调用gpt记进去了。 1 个帖子 - 1 位参与者 阅读完整话题

geminiflash(反代)如何提升指令遵循度

各位佬,gemini如何提升指令遵循度?比如我让他写100条,他总是会自己缩水。加了约束提示词依旧不能完成任务。3.7flash和3.8flash同样有这样的问题,虽然TPS很快,但是智商太低了。这种情况该怎么解决?还是说agy中才能发挥出gemini的最大性能? 3 个帖子 - 3 位参与者 阅读完整话题

现在有一台闲置坏电脑怎么处理

手里有台拯救者r9000p2022锐龙款,应该是批次问题导致的缩肛,重植过一次,然后现在的情况是整个电脑没法开机,换主板的话应该就完全能恢复,但是价格有点高,但是砸手里又不知道能干嘛,来问问有没有有类似经验的佬 7 个帖子 - 5 位参与者 阅读完整话题