# 05｜信息源与核验说明

## 说明
这份 deep research 重点参考了 2025–2026 年的：

- 模型厂商官方发布
- 官方开发文档
- 大型科技公司研究/趋势报告
- 一线投资机构/创业机构对于应用层机会的判断

我优先采用了**直接可核验来源**，并在浏览器工具不可用的情况下，通过可访问的文本抓取方式对关键页面进行了二次核验。

## 核验时间
- 核验时间：`2026-07-09 14:06:35 UTC`

## 直接核验的核心来源

### 1. OpenAI｜New tools for building agents
- 链接：<https://openai.com/index/new-tools-for-building-agents/>
- 核心核验点：
  - Responses API
  - built-in tools: web search / file search / computer use
  - Agents SDK
  - observability tools
- 研究中用途：用于判断 AI 已从“回答”进入“可执行工作流”的平台能力阶段。

### 2. OpenAI｜The state of enterprise AI (2025 report)
- 链接：<https://openai.com/business/guides-and-resources/the-state-of-enterprise-ai-2025-report/>
- 核心核验点：
  - 超过 1 million business customers 使用 OpenAI tools
  - ChatGPT message volume grew 8x
  - API reasoning token consumption per organization increased 320x YoY
  - Enterprise users report saving 40–60 minutes per day
  - median sector grew by more than 6x, tech sector 11x
- 研究中用途：用于验证企业真实采用正在加速，而不是只有消费者层面热闹。

### 3. Anthropic｜Introducing Claude 4
- 链接：<https://www.anthropic.com/news/claude-4>
- 核心核验点：
  - Claude 4 强调 coding / advanced reasoning / AI agents
  - extended thinking with tool use
  - tools in parallel
  - improved memory capabilities
  - Claude Code GA
- 研究中用途：用于判断 coding agents、tool use、长任务协作已进入高可用阶段。

### 4. Anthropic｜Economic Index: Tracking AI's role in the economy
- 链接：<https://www.anthropic.com/research/economic-index-geography>
- 核心核验点：
  - software engineering is still by far in the lead
  - directively automated tasks from 27% → 39%
  - API users are significantly more likely to automate tasks
- 研究中用途：用于判断真实世界中 AI 使用正从辅助走向自动化，且工程/软件仍是领先场景。

### 5. Google｜Gemini 2.5: Our most intelligent AI model
- 链接：<https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/>
- 核心核验点：
  - Gemini 2.5 is a thinking model
  - strong reasoning and code capabilities
  - context-aware agents
- 研究中用途：用于确认“thinking / reasoning / code”已经是主流模型竞争重点。

### 6. GitHub Copilot documentation
- 链接：<https://docs.github.com/en/copilot>
- 核心核验点：
  - cloud agent
  - automations
  - autonomous task completion
  - parallel task execution
- 研究中用途：用于支持“编码代理已产品化并进入结构化能力阶段”的判断。

### 7. Microsoft｜2025 Work Trend Index: The Frontier Firm is born
- 链接：<https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born>
- 核心核验点：
  - 31,000 workers across 31 countries
  - 82% leaders: pivotal year to rethink strategy/operations
  - 81% expect agents moderately/extensively integrated in 12–18 months
  - 24% already deployed AI organization-wide
  - phase 2/3 描述：digital colleagues → agents running entire workflows
- 研究中用途：用于支持“agent 将重塑组织工作方式，而不仅是个人助手”的判断。

### 8. Sequoia｜AI 50: AI Agents Move Beyond Chat
- 链接：<https://www.sequoiacap.com/article/ai-50-2025/>
- 核心核验点：
  - AI move from responding to prompts to solving problems and completing workflows
  - application layer tools create real business results
  - 典型案例：Harvey / Sierra / Cursor
- 研究中用途：用于支持“应用层机会从 copilot 转向 workflow completion”的核心结论。

### 9. Y Combinator｜Requests for Startups 2025
- 链接：<https://www.ycombinator.com/rfs?year=2025>
- 核心核验点：
  - AI models are improving really fast
  - now able to do complex work far beyond engineering
  - next step: AI-native companies that don’t sell software—they sell the service
- 研究中用途：用于支持“AI 原生服务化业务”这一高优先级方向。

### 10. a16z｜How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025
- 链接：<https://a16z.com/ai-enterprise-2025/>
- 核心核验点：
  - AI budgets expected average ~75% growth next year
  - 37% respondents use 5 or more models
  - enterprise model layer has not become commoditized
  - Anthropic 在 engineering/coding 场景强势
  - Gemini 2.5 Flash price/performance highlighted
- 研究中用途：用于支持“企业需求真实增长、multi-model 成为常态、应用选择按场景分化”的判断。

### 11. a16z｜AI Voice Agents: 2025 Update
- 链接：<https://a16z.com/ai-voice-agents-2025-update/>
- 核心核验点：
  - lower latency and improved performance
  - voice agents market exploded in H2 2024
  - 22% of the most recent YC class represented voice-building companies（转引 Cartesia 数据）
  - 应从 wedge 切入，而不是一上来全替代人工
  - 高 BPO / call center spend 场景更适合切入
- 研究中用途：用于支持语音代理机会的边界与切入策略。

## 研究限制与说明

### 1. 浏览器交互工具本轮不可用
本轮环境中的 browser 工具缺少底层可执行组件，因此未采用浏览器快照式核验；改用终端 + 文本抓取方式完成二次验证。

### 2. 个别页面对抓取不友好
例如 Stripe 相关页面在文本代理下出现了较多导航噪音，未作为本研究的主要定量依据，仅作为辅助观察来源，不进入核心论证链。

### 3. 投资机构观点不等于客观事实
a16z、Sequoia、YC、BVP 等来源适合帮助判断“机会方向”和“创业热点”，但它们天然带有投资立场。因此本研究中，这类来源主要用于：

- 识别机会聚焦点
- 识别产品/市场演化方向
- 辅助排序

而不是单独作为事实证据。

## 本研究的证据结构
整份研究的判断，主要建立在 3 层证据上：

1. **官方能力层证据**  
   OpenAI / Anthropic / Google / GitHub 官方发布与文档
2. **企业采用层证据**  
   OpenAI enterprise report / Microsoft Work Trend Index / Anthropic Economic Index
3. **创业与应用层证据**  
   Sequoia / YC / a16z 对应用落地方向的观察

## 最终说明
这份研究适合拿来做：

- 方向判断
- 机会筛选
- 30 天验证选题
- 产品/服务切入思考

如果你下一步要做更具体的决策，例如：

- “我该做哪个具体赛道？”
- “我该从哪个用户群开始？”
- “我该先做服务还是先做产品？”

那最合理的下一步，是基于这份总报告，再做一轮：

> **面向你本人背景和资源约束的二次聚焦研究。**