05|信息源与核验说明¶
说明¶
这份 deep research 重点参考了 2025–2026 年的:
- 模型厂商官方发布
- 官方开发文档
- 大型科技公司研究/趋势报告
- 一线投资机构/创业机构对于应用层机会的判断
我优先采用了直接可核验来源,并在浏览器工具不可用的情况下,通过可访问的文本抓取方式对关键页面进行了二次核验。
核验时间¶
- 核验时间:
2026-07-09 14:06:35 UTC
直接核验的核心来源¶
1. OpenAI|New tools for building agents¶
- 链接:https://openai.com/index/new-tools-for-building-agents/
- 核心核验点:
- Responses API
- built-in tools: web search / file search / computer use
- Agents SDK
- observability tools
- 研究中用途:用于判断 AI 已从“回答”进入“可执行工作流”的平台能力阶段。
2. OpenAI|The state of enterprise AI (2025 report)¶
- 链接:https://openai.com/business/guides-and-resources/the-state-of-enterprise-ai-2025-report/
- 核心核验点:
- 超过 1 million business customers 使用 OpenAI tools
- ChatGPT message volume grew 8x
- API reasoning token consumption per organization increased 320x YoY
- Enterprise users report saving 40–60 minutes per day
- median sector grew by more than 6x, tech sector 11x
- 研究中用途:用于验证企业真实采用正在加速,而不是只有消费者层面热闹。
3. Anthropic|Introducing Claude 4¶
- 链接:https://www.anthropic.com/news/claude-4
- 核心核验点:
- Claude 4 强调 coding / advanced reasoning / AI agents
- extended thinking with tool use
- tools in parallel
- improved memory capabilities
- Claude Code GA
- 研究中用途:用于判断 coding agents、tool use、长任务协作已进入高可用阶段。
4. Anthropic|Economic Index: Tracking AI's role in the economy¶
- 链接:https://www.anthropic.com/research/economic-index-geography
- 核心核验点:
- software engineering is still by far in the lead
- directively automated tasks from 27% → 39%
- API users are significantly more likely to automate tasks
- 研究中用途:用于判断真实世界中 AI 使用正从辅助走向自动化,且工程/软件仍是领先场景。
5. Google|Gemini 2.5: Our most intelligent AI model¶
- 链接:https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/
- 核心核验点:
- Gemini 2.5 is a thinking model
- strong reasoning and code capabilities
- context-aware agents
- 研究中用途:用于确认“thinking / reasoning / code”已经是主流模型竞争重点。
6. GitHub Copilot documentation¶
- 链接:https://docs.github.com/en/copilot
- 核心核验点:
- cloud agent
- automations
- autonomous task completion
- parallel task execution
- 研究中用途:用于支持“编码代理已产品化并进入结构化能力阶段”的判断。
7. Microsoft|2025 Work Trend Index: The Frontier Firm is born¶
- 链接:https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born
- 核心核验点:
- 31,000 workers across 31 countries
- 82% leaders: pivotal year to rethink strategy/operations
- 81% expect agents moderately/extensively integrated in 12–18 months
- 24% already deployed AI organization-wide
- phase 2/3 描述:digital colleagues → agents running entire workflows
- 研究中用途:用于支持“agent 将重塑组织工作方式,而不仅是个人助手”的判断。
8. Sequoia|AI 50: AI Agents Move Beyond Chat¶
- 链接:https://www.sequoiacap.com/article/ai-50-2025/
- 核心核验点:
- AI move from responding to prompts to solving problems and completing workflows
- application layer tools create real business results
- 典型案例:Harvey / Sierra / Cursor
- 研究中用途:用于支持“应用层机会从 copilot 转向 workflow completion”的核心结论。
9. Y Combinator|Requests for Startups 2025¶
- 链接:https://www.ycombinator.com/rfs?year=2025
- 核心核验点:
- AI models are improving really fast
- now able to do complex work far beyond engineering
- next step: AI-native companies that don’t sell software—they sell the service
- 研究中用途:用于支持“AI 原生服务化业务”这一高优先级方向。
10. a16z|How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025¶
- 链接:https://a16z.com/ai-enterprise-2025/
- 核心核验点:
- AI budgets expected average ~75% growth next year
- 37% respondents use 5 or more models
- enterprise model layer has not become commoditized
- Anthropic 在 engineering/coding 场景强势
- Gemini 2.5 Flash price/performance highlighted
- 研究中用途:用于支持“企业需求真实增长、multi-model 成为常态、应用选择按场景分化”的判断。
11. a16z|AI Voice Agents: 2025 Update¶
- 链接:https://a16z.com/ai-voice-agents-2025-update/
- 核心核验点:
- lower latency and improved performance
- voice agents market exploded in H2 2024
- 22% of the most recent YC class represented voice-building companies(转引 Cartesia 数据)
- 应从 wedge 切入,而不是一上来全替代人工
- 高 BPO / call center spend 场景更适合切入
- 研究中用途:用于支持语音代理机会的边界与切入策略。
研究限制与说明¶
1. 浏览器交互工具本轮不可用¶
本轮环境中的 browser 工具缺少底层可执行组件,因此未采用浏览器快照式核验;改用终端 + 文本抓取方式完成二次验证。
2. 个别页面对抓取不友好¶
例如 Stripe 相关页面在文本代理下出现了较多导航噪音,未作为本研究的主要定量依据,仅作为辅助观察来源,不进入核心论证链。
3. 投资机构观点不等于客观事实¶
a16z、Sequoia、YC、BVP 等来源适合帮助判断“机会方向”和“创业热点”,但它们天然带有投资立场。因此本研究中,这类来源主要用于:
- 识别机会聚焦点
- 识别产品/市场演化方向
- 辅助排序
而不是单独作为事实证据。
本研究的证据结构¶
整份研究的判断,主要建立在 3 层证据上:
- 官方能力层证据
OpenAI / Anthropic / Google / GitHub 官方发布与文档 - 企业采用层证据
OpenAI enterprise report / Microsoft Work Trend Index / Anthropic Economic Index - 创业与应用层证据
Sequoia / YC / a16z 对应用落地方向的观察
最终说明¶
这份研究适合拿来做:
- 方向判断
- 机会筛选
- 30 天验证选题
- 产品/服务切入思考
如果你下一步要做更具体的决策,例如:
- “我该做哪个具体赛道?”
- “我该从哪个用户群开始?”
- “我该先做服务还是先做产品?”
那最合理的下一步,是基于这份总报告,再做一轮:
面向你本人背景和资源约束的二次聚焦研究。