DeepSeek’s compression of active key-value memory attacks a real cost center of long-context agents. But shrinking memory per token does not remove the difficulty of operating its very large composite model.
01 · Tech News
Follow the AI industry.
74 reports following 43 companies and partnerships from model release to physical deployment.
All Tech News
Browse every report, including the current lead.
5 articles · By event date
OpenAI's GPT-6 Astra can finish some agent tasks in fewer steps while raising the stakes of unsupervised execution. Efficiency is part of the case for Astra; it is not a reason to give the model more authority.
Google kept the introductory token rate identical to its predecessor, but one independent test found that higher reasoning effort and extra agent loops pushed the cost of finished work up by roughly 40 percent.
Alibaba’s Qwen3.8 puts a 2.4-trillion-parameter model in public repositories. That is a real expansion of freedom—but mainly for organizations already equipped to operate industrial AI infrastructure.
Enterprise buyers are being asked to treat multimodal tool use and coding scores as evidence of business output. But moving from a benchmark score to economic return depends on the unglamorous friction of permissions, harnesses, and human review.




