黄超老师与 HKUDS 全量尽调报告
快照时间:2026-08-12 07:16:37 CST(UTC+08:00) 报告范围:PI、团队、研究演进、论文、开源项目、组织模式、招生路径、 Pengyi 适配度、风险与下一步切入口。 证据规则:事实只来自官方页面、论文页和公开仓库;stars、引用数和奖项若 来自团队自述则明确标注;研究判断与组织判断均标记为推断。
0. 一页结论
总判断
黄超老师组是 Pengyi 当前 P0 级博士、RA/长期远程研究合作目标。真正的 匹配点不是笼统的“大家都在做 Agent”,而是双方都在尝试跑通:
研究问题
-> 可证伪方法与评估
-> 可运行开源系统
-> 用户、贡献者与真实 failure cases
-> 新研究问题
-> 下一轮论文和系统迭代
对 HKUDS 更准确的定义是:
以图学习、推荐、数据挖掘和时空智能为学术底座,向上建立 RAG、Agent runtime、tool/skill/memory/harness 和 multi-Agent 能力,再将这些能力部署 到科研、编程、教育、交易、视频、软件控制和个人生产力场景中,并通过 开源社区形成研究与工程反馈。
立即决策
| 项目 | 判断 |
|---|---|
| 研究方向匹配 | 高 |
| 开源工程匹配 | 很高,已有 4 个 merged PR |
| 金融场景桥梁 | 高,Vibe-Trading / AI-Trader 可承接 |
| 当前论文证据 | 弱,尚无冻结的顶会论文成果 |
| 方法与统计严谨度 | 中低,必须用 benchmark 和 protocol 补齐 |
| 经费与名额确定性 | 未知,不得把“招收”理解为已资助 |
| 导师日常指导模式 | 未知,必须通过交流和组员访谈核实 |
| 当前动作 | 正式申请、证据型交流、持续 OSS 三线并行 |
推荐主切入口
Leakage-Safe and Failure-Aware Evaluation for Agentic Quant Research Systems,即“面向 Agentic 量化研究系统的防泄漏与故障感知评估”。
它把 HKUDS 的 DeepResearch evaluation、OpenHarness、LightRAG 和 Vibe-Trading 与 Pengyi 的 Quant/FICC、PAT、PPRT 真实证据连接起来,而且 能形成论文、benchmark、harness 和开源仓库四类成果。所有效果目前均为 UNMEASURED。
1. 证据口径与限制
本报告按四类证据分层:
- 官方事实:HKU 个人主页、实验室主页、招生页、HKU CDS 招生页。
- 论文事实:论文官网、arXiv、正式 venue 信息。
- 开源快照:GitHub organization、repository、release 和 PR 页面。
- 分析推断:从研究与项目组合推导出的架构、组织方式和适配判断。
以下内容不能直接当成独立验证的研究质量:
- stars、forks、GitHub Trending 次数;
- 实验室或个人主页自述的引用量、h-index 和奖项;
- README 中未经独立复现的性能或能力声明;
- 项目数量和更新速度。
指标还存在明显快照差异:HKU 个人页写有超过 14,000 次引用和约 77,000 GitHub stars,而实验室 publications 页显示 19,000+ citations、h-index 70, 本次 GitHub API 聚合则约为 350,921 stars。这些数字很可能来自不同更新时间、 统计范围或口径,故本报告只把它们作为“公开影响力自述/分发规模”,不汇总成 单一确定指标。
2. PI 全景
2.1 已核实身份
HKU 官方页面显示,Chao Huang(黄超)是香港大学数据科学研究院与计算与 数据科学学院的 Assistant Professor、PhD supervisor,并领导 Data Intelligence Lab。当前公开研究重点包括:
- Large Language Models;
- AI Agents;
- Graph Machine Learning。
其公开经历包括 University of Notre Dame 博士和 JD Research America Scientist。HKU 官方个人页还列出 WAIC 2024 Bright Star、2024 Frontier Science Award、2025 AI100 Young Pioneers、2025 Global AI 2000 等荣誉; 这些是官方个人页展示的信息,本报告未逐项进行奖项主办方的二次核验。
2.2 研究者画像
事实:研究跨越图学习、推荐、时空智能、RAG、AI Agents 和开源系统。
推断:黄超老师并非只做“模型论文”或只做“开源产品”,而是在推动一种 paper-system-community 联动的研究生产方式。对于组员,这很可能意味着:
- 需要同时具备研究定义、工程落地和公开沟通能力;
- 项目必须较快形成 runnable artifact;
- 研究价值、维护成本与社区反馈会同时进入项目决策;
- 学生可能获得较强的项目 ownership,也会承担较高的交付和维护压力。
上述组织判断需要通过老师和组员交流核实,不能仅凭仓库推断。
3. 团队与项目责任图
实验室官方 Group & Join Us 页面在本次快照中列出 20 位 Group Members。 页面未对所有人统一标明 PhD、MPhil、RA 或 visitor 身份,因此这里只记录名字与 公开项目关联,不擅自推断学籍和雇佣关系。
| 成员 | 官方页可见的代表项目/论文关联 | 可用于理解的能力节点 |
|---|---|---|
| Lianghao Xia | AI-Researcher;AnyGraph、OpenGraph 等 | 自动科研、图基础模型 |
| Yuhao Yang | CLI-Anything、ClawWork;GraphAgent | 工具化、Agent 与图 |
| Yangqin Jiang | OpenPhone;RecLM、RecGPT | 移动 Agent、推荐 |
| Xubin Ren | nanobot、VideoRAG;RLMRec | Agent runtime、多模态 RAG |
| Jiabin Tang | OpenHarness、AutoAgent、AI-Researcher、ClawTeam;GraphGPT、HiGPT | Harness、多 Agent、科研 Agent |
| Zhonghang Li | DeepCode、FastCode;UrbanGPT、FlashST、GPT-ST | Coding Agent、时空智能 |
| Zongwei Li | DeepCode、FastCode;推荐研究 | Coding Agent、推荐 |
| Guoxuan Chen | SepLLM | 高效 LLM |
| Tianyu Fan | AI-Trader、MiniRAG;DeepResearch evaluation | 交易 Agent、轻量 RAG、评估 |
| Zirui Guo | LightRAG、RAG-Anything | Graph-RAG、多模态 RAG |
| Lingrui Xu | OpenSpace、Paper2Slides、VideoRAG | Skill 生态、科研传播、多模态 |
| Jingyuan Wang | LightReasoner | 轻量推理 |
| Lingxuan Huang | ViMax | 视频 Agent |
| Yuxuan Chen | OpenSpace | Agent skill 与演进 |
| Bingxi Zhao | DeepTutor | 教育 Agent |
| Haozhe Wu | Vibe-Trading | Agentic finance research |
| Junzhou Fang | 官方页本次未列明确项目 | 待核实 |
| Jinyuan Chen | 官方页本次未列明确项目 | 待核实 |
| Yumeng Wang | 官方页本次未列明确项目 | 待核实 |
| Jiahao Zhang | 官方页本次未列明确项目 | 待核实 |
3.1 对 Pengyi 最相关的技术节点
按当前项目证据,不代表联系优先级:
- Jiabin Tang:OpenHarness / AutoAgent / AI-Researcher / ClawTeam;
- Zirui Guo:LightRAG / RAG-Anything;
- Tianyu Fan:AI-Trader / MiniRAG / DeepResearch evaluation;
- Haozhe Wu:Vibe-Trading。
正式申请和首次研究联系仍应遵循官方 join 页面,直接面向黄超老师并使用规定 的邮件主题,不做无差别组员群发。
4. 研究演进:2021–2026
4.1 2021–2023:图、推荐与时空学习基础
这一阶段的核心积累包括 graph contrastive learning、self-supervised recommendation、knowledge graph recommendation、spatio-temporal forecasting。它形成了后续系统所需的结构化表示、关系推理、排序和时空建模 能力。
4.2 2023–2024:图基础模型与 LLM 接入
GraphGPT、HiGPT、OpenGraph、AnyGraph、UrbanGPT、LLMRec、RecLM 等工作 表明研究问题从“单任务模型”扩展到:
- 图上的通用表示和 foundation model;
- LLM 与图、推荐和城市数据的连接;
- 更统一的任务接口和跨数据集迁移。
4.3 2024–2025:RAG 与 Agentic research
LightRAG、MiniRAG、VideoRAG、RAG-Anything 将研究中心推向异构知识检索、 图结构上下文、轻量模型与多模态。AutoAgent、AI-Researcher 和 DeepResearch-Eval 则开始处理 Agent 的构建、科研流程与可测量性。
4.4 2025–2026:Agent 全栈与垂直应用
公开组合已覆盖:
Model / Provider
-> Retrieval and Context
-> Tool / Skill / Memory
-> Harness / Permission / Session
-> Multi-Agent Coordination
-> Domain Workflow
-> Evaluation and Human Review
CLI-Anything、nanobot、OpenHarness、OpenSpace、ClawTeam 等更接近平台层; DeepTutor、Vibe-Trading、AI-Trader、OpenPhone、ViMax 等则验证领域适配。
4.5 研究方向的连续性
“HKUDS 已从图学习切换成 Agent 应用”这一说法并不准确。更合理的解释是:
图、推荐、时空和数据智能仍是底座;RAG 与 Agents 成为新的系统组织方式; 领域项目是验证和扩散渠道。
5. HKUDS 总系统架构
Layer A Data Intelligence Foundations
Graph / Recommendation / Spatio-temporal / Ranking
|
Layer B Context and Retrieval
LightRAG / MiniRAG / RAG-Anything / VideoRAG
|
Layer C Agent Runtime and Capabilities
Tools / Skills / Memory / Harness / Permission / Multi-Agent
|
Layer D Research and Coding Systems
AI-Researcher / DeepResearch-Eval / DeepCode / Paper2Slides
|
Layer E Vertical Applications
Education / Trading / Mobile / Video / Smart City
|
Layer F Open-source Feedback and Distribution
Users / Contributors / Issues / Releases / Failure Cases
这不是 HKUDS 发布的正式组织图,而是本报告对公开研究与代码组合的架构推断。
6. 论文矩阵
6.1 2026 年代表工作
| 论文 | Venue | 核心问题 | 对 Pengyi 的价值 |
|---|---|---|---|
| VideoRAG | KDD 2026 | 长视频的文本图与多模态检索 | 多模态 context/eval |
| AutoAgent | ACL 2026 | 以自然语言构建与管理 Agent | Agent OS 与 harness |
| MiniRAG | ACL 2026 | 小模型、异构图与拓扑增强检索 | 成本受限 RAG |
| OpenPhone | ACL 2026 | 移动终端 Agent | tool/action boundary |
| AnyGraph | ACL 2026 | 跨图任务的基础模型 | 通用图表示 |
6.2 2025 年代表工作
| 论文 | Venue | 核心问题 | 对 Pengyi 的价值 |
|---|---|---|---|
| LightRAG | EMNLP 2025 | 图与向量的双层检索、增量更新 | 已有 3 个 merged PR |
| AI-Researcher | NeurIPS 2025 | 自动科研全流程 | PPHD / PRDT 直接相关 |
| Justice or Prejudice? | ICLR 2025 | LLM-as-a-Judge 偏差 | Agent evaluation |
| Aria-UI | ACL 2025 | GUI/UI Agent | action evaluation |
| SepLLM | ICML 2025 | 高效序列建模 | efficiency thinking |
6.3 最应先读的四篇
- Understanding DeepResearch via Reports:报告质量、冗余、事实性、
LLM judge 与专家一致性;100 个 query、12 类任务、4 个商业系统。
- AutoAgent:Agent 构建、系统工具、可执行引擎与自定义。
- LightRAG:图增强检索、双层 retrieval 与增量更新。
- AI-Researcher:从文献/想法到算法、实验验证与论文写作的自动科研链路。
这四篇分别对应 Evaluation -> Harness -> Retrieval -> Research Workflow, 足以支撑第一次高质量交流。PPHD Paper Reviewer 应对每篇记录 claim、baseline、 数据、metric、reproducibility、failure 和 extension,而非只写摘要。
7. 开源组合快照
7.1 GitHub API 快照
快照时间:2026-08-12 07:14:34 CST(UTC+08:00)。
| 指标 | 数值 |
|---|---|
| Public repositories | 91 |
| Aggregate stars | 350,921 |
| Aggregate forks | 48,358 |
| 2026 年以来有 push 的仓库 | 34 |
| 最近 90 天有 push 的仓库 | 23 |
| Archived repositories | 0 |
| Python 为 primary language | 78 |
| TypeScript 为 primary language | 3 |
| Jupyter Notebook 为 primary language | 2 |
这些是时间敏感的分发与活跃度快照,不是论文质量或维护质量结论。
7.2 Stars 最高的代表项目
| Repository | Stars at snapshot | 架构位置 |
|---|---|---|
| CLI-Anything | 46,905 | Tool / software control |
| nanobot | 46,858 | Agent runtime |
| LightRAG | 38,778 | Retrieval |
| DeepTutor | 34,668 | Education Agent |
| Vibe-Trading | 30,626 | Finance vertical |
| RAG-Anything | 22,874 | Multimodal RAG |
| AI-Trader | 21,247 | Trading Agent |
| DeepCode | 16,333 | Coding / research artifact |
| OpenHarness | 15,317 | Agent harness |
| ViMax | 11,839 | Video Agent |
| AutoAgent | 9,734 | Agent construction |
| ClawWork | 8,336 | Agent work organization |
| OpenSpace | 7,361 | Skill retrieval/evolution |
| AI-Researcher | 5,663 | Scientific Agent |
| ClawTeam | 5,484 | Multi-Agent team |
7.3 对开源组合的判断
优势:项目可运行、传播强、覆盖 runtime 到 vertical application、能获得 真实用户 failure cases,并且存在 paper-code linkage。
风险:项目数量和传播速度可能带来维护负担;stars 会奖励易传播产品而非 严谨研究;不同仓库的数据、benchmark、测试成熟度可能差异很大;多个相似 Agent 项目需要核实复用边界和长期 owner。
8. 组织与研究生产模式
8.1 可观察事实
- 学生名字与多个公开项目/论文绑定;
- 仓库覆盖 paper implementation、framework、product 和 benchmark;
- 多个项目持续 release,并接受外部 issue/PR;
- HKU IDS 2025 Research Seed Funds 公告列出黄超老师为
Towards Autonomous Scientific Research with LLM Agents 的 PI,研究主题 包括 AI for Science、LLMs、AI Agents。
8.2 推断的运行模型
PI research portfolio
-> student/project ownership
-> paper and open-source artifact
-> community adoption and failures
-> benchmark/method refinement
-> new project or paper
这套模式对能主动定义问题、稳定交付、维护开源项目的人有吸引力;对需要高度 结构化手把手指导、难以承受公开交付压力或无法长期聚焦的人可能不理想。
8.3 必须面谈核实的问题
- 每周与导师/直接 mentor 的交流频率和形式是什么?
- 课题由 PI 指定、学生提出还是项目演化产生?
- paper、repo、产品化和维护分别如何计入学生优先级?
- authorship、project ownership 和长期 maintenance 如何约定?
- 计算资源、私有数据和实验预算如何分配?
- 组内如何执行 ablation、统计检验、negative result 和复现?
- 自费名额、奖学金名额、RA 和 remote intern 的具体边界是什么?
9. 招生、合作与时间线
9.1 实验室公开入口
官方 join 页面显示:
- 持续招收 PhD/MPhil applicants;
- PhD/MPhil 申请人在正式申请中选择黄超老师,并可 case-by-case 讨论;
- 页面提到若干 self-financed PhD/MPhil positions;
- 欢迎 RA、remote intern 和 visitor,优先希望持续至少 6 个月;
- 邮件主题应为
Prospective Student: Your Name - Your Affiliation; - 邮件需概述教育、研究成果、编程/理论能力,并附 PDF resume。
self-financed 不等于已获得 scholarship,也不代表任何个人已有录取或经费。
9.2 HKU CDS 2027 入学节点
| Route | Official deadline/status | PPHD internal action |
|---|---|---|
| Second Early Round | 2026-08-31 23:59 | 2026-08-26 前冻结完整包 |
| Main Round | 2026-09-01 至 2026-12-01 | 2026-11-20 前完成获批提交 |
| HKPFS | 通常 9 月至 12 月初;2027 精确 RGC 节点需复核 | 2026-09-01 重查官方页 |
HKU CDS 页面列出的 English requirement 包括在不满足英语授课学历豁免时提供 TOEFL iBT 85+ 等合格证明,并要求 research proposal。支持材料页面还要求两位 referees;PhD applicant 的论文/英文研究写作样本等具体要求应按申请系统再次 核对。
9.3 建议三线并行
- Formal application:不把交流结果当作正式录取替代品。
- Evidence-led conversation:以研究问题和 4 个 merged PR 为入口。
- Continued OSS contribution:只做有边界、可测试、有价值的修复。
10. Pengyi 已有证据
截至本报告时间,GitHub 公开页面核实了 4 个由 pengpengyi92 提交并合并的 HKUDS PR:
| Repository / PR | Merge time (UTC) | 保护的边界或 invariant |
|---|---|---|
| LightRAG #3574 | 2026-08-05 10:38:12 | 在 query boundary 验证 rerank result |
| LightRAG #3607 | 2026-08-10 10:12:29 | 暴露 token-limit truncation,拒绝空截断结果 |
| LightRAG #3622 | 2026-08-11 13:40:12 | entity rename 时保留 self-loop |
| Vibe-Trading #1066 | 2026-08-11 20:00:08 | 正确处理零波动率下的 discounted forward value |
这些 PR 能证明:
- 能读懂真实上游项目;
- 能定位边界条件和系统 invariant;
- 能做 scoped fix、测试和 maintainer 协作;
- 已在 RAG 与 finance 两个相关项目建立可验证互动。
它们不能证明:已达到博士研究方法要求、已有论文能力、已获得导师认可或录取。
11. Pengyi 适配度矩阵
| 维度 | 当前评级 | 证据 | 主要缺口 |
|---|---|---|---|
| Open-source engineering | 很高 | 4 merged PR、PPRT 流程 | 需保持持续质量 |
| Agent architecture | 高 | PAT、Pcaml、PRDT 等系统设计 | 很多仍是计划,缺 measured result |
| RAG reliability | 高 | 3 个 LightRAG merged fixes | 缺冻结 benchmark/论文问题 |
| Finance domain | 高 | Quant/FICC/交易研究背景、Vibe PR | 缺 PIT 数据实验与现实成本模型 |
| Research formulation | 中 | 已提出 falsifiable direction | 问题仍需进一步收窄 |
| Experimental rigor | 中低 | 有 benchmark 意识 | 缺数据、protocol、统计检验和 ablation |
| Publication evidence | 低 | 暂无冻结的顶会成果 | 需形成首篇可复现研究 |
| Mathematical depth | 待证明 | 简历/项目不能替代现场证据 | 需显式公式、推导和实验设计 |
| English research writing | 待证明 | 有英文工程沟通 | 需 1–2 页研究 pitch 和 sample |
| Long-horizon focus | 风险项 | 项目丰富 | repo 过多、context switching |
11.1 真实优势
Pengyi 的差异化不是“有很多 Agent repo”,而是:
- 金融领域约束、数据时间性、风险与数值边界意识;
- 能从 issue 进入上游代码,保护 invariant 并完成合并;
- 能把系统架构、evaluation 和开源交付放到同一条链上;
- 有机会把 HKUDS 的通用 Agent stack 推入一个严格、真实、可复现的金融场景。
11.2 最大风险
最大的风险不是不会做项目,而是 项目过多导致研究问题不够深。申请材料中 不得展示完整 Pengyi OS 目录来替代研究成果。应该只展示:
one research question
+ one benchmark protocol
+ four merged PRs
+ one reproducible pilot
+ one honest limitation section
12. 主研究切入口
12.1 研究问题
在固定 point-in-time 数据、工具、模型预算和市场时间戳条件下,RAG、memory、 harness 与 multi-Agent coordination 中哪些能力真正提升量化研究报告的事实 支持、可复现性和故障恢复,哪些只会增加成本、方差或 unsupported claims?
12.2 与 HKUDS 资产的连接
| HKUDS asset | 可承接问题 |
|---|---|
| Understanding DeepResearch via Reports | 报告质量、事实性、judge 设计 |
| OpenHarness / AutoAgent | agent loop、session、tool、permission、recovery |
| LightRAG / MiniRAG | retrieval、graph context、small-model cost |
| Vibe-Trading / AI-Trader | 金融任务、数据时点、策略和报告 |
| AI-Researcher | hypothesis、implementation、validation、writing workflow |
12.3 冻结实验设计
Public or synthetic point-in-time finance tasks
|
+-> deterministic non-Agent baseline
+-> single-Agent baseline
+-> Agent + RAG
+-> Agent + memory
+-> multi-Agent / harness
+-> stale-data and tool-failure variants
|
-> repeated runs + replay + human audit
建议指标:
- factual support / citation correctness;
- temporal leakage rate;
- unsupported claim rate;
- result variance across runs;
- tool failure detection and recovery;
- replay success and reproducibility;
- latency、token/API cost、human correction time;
- 简单金融结果的一致性与风险约束违反率。
所有 baseline 结果、效果提升和成本数据均为 UNMEASURED,直到任务集、数据 版本、模型版本、seed、预算和 stop rule 固定并真正运行。
12.4 可发表性判断
只有同时满足以下条件,才值得向论文推进:
- 新 benchmark 暴露现有 Agent evaluation 无法观察的 failure;
- 提供可重复的 point-in-time task/data protocol;
- 至少有一个方法或 harness 改动在冻结数据上显著改善核心 metric;
- 有 negative result、ablation 和跨模型/跨任务稳健性;
- 金融 domain 不是换皮,而是带来时间泄漏、成本、风险和执行约束。
13. 备选切入口
Route B:Dynamic Graph-RAG Reliability
围绕 LightRAG 已有贡献研究动态语料下的 graph update invariant、retrieval corruption、truncation、stale state 与 replay。优势是 maintainer evidence 最强; 不足是金融差异化较弱。
Route C:Agent Harness Recovery and Governance
围绕 OpenHarness 研究 permission、session recovery、tool failure、trace 和 multi-Agent coordination benchmark。优势是系统性强;不足是需要先建立直接 贡献和严谨任务集。
Route D:Scientific-Agent Audit
连接 AI-Researcher 与 DeepResearch-Eval,审计 literature evidence、idea novelty、implementation correctness、experiment validity 和 report factuality。 优势是与 PPHD/PRDT 强相关;不足是竞争激烈,必须有明确新方法或 benchmark。
14. 联系前的硬门槛
以下内容未完成前,不发送最终邮件:
- [ ] 一页研究 pitch 已冻结;
- [ ] 四个 merged PR evidence card 已冻结;
- [ ] DeepResearch ReportEval、AutoAgent、LightRAG 三篇 hostile review 完成;
- [ ] 一段真实 mismatch/limitation 已写入;
- [ ] Research CV 和 PDF 体积已检查;
- [ ] 邮件符合官方主题格式且不超过约 180 英文词;
- [ ] funding、admission、paper readiness 无夸大;
- [ ] Human Send Gate 已明确批准。
推荐邮件主题:
Prospective Student: Pengyi Peng - [Current Affiliation]
推荐 opening 的事实骨架:
I recently contributed four merged fixes across LightRAG and Vibe-Trading while studying failure boundaries in retrieval and financial research systems. I am now formulating a leakage-safe, replayable evaluation protocol for agentic quantitative research, connecting report factuality, harness recovery and point-in-time market data.
这只是报告中的草案,不构成已批准邮件。
15. 72 小时执行图
T+12h
- 完成四 PR evidence card;
- 对
Understanding DeepResearch via Reports完成第一份 hostile review; - 把主问题压缩为一页,所有结果标记
UNMEASURED。
T+36h
- 完成 AutoAgent、LightRAG fit card;
- 冻结 benchmark task schema、baseline、metrics、budget 和 stop rule;
- 完成 CV + research evidence index。
T+72h
- 做 PPHD hostile review;
- 输出
SEND / REVISE / HOLD; - 若
SEND,发出一次精准联系; - 同时继续 HKU early round 正式材料,不等待邮件回复。
16. 30/60/90 天路线
| 时间 | 必须形成的证据 | Stop rule |
|---|---|---|
| 30 天 | frozen tasks、deterministic baseline、single-Agent baseline、failure taxonomy | 无可复现 task 则停止扩系统 |
| 60 天 | RAG/memory/harness ablation、重复运行、初步统计、公开 repo | 没有稳定差异则收缩问题 |
| 90 天 | technical report、完整 benchmark card、targeted paper plan 或明确 negative result | 不为追热点伪造正结果 |
17. 最终风险清单
学术风险
- 研究问题被大型 Agent benchmark 覆盖,缺乏新颖性;
- 金融任务数据不可公开或容易泄漏;
- LLM judge 偏差导致评估不可信;
- 成果更像 engineering benchmark 而非 method paper。
组织风险
- PI 与学生项目过多,实际指导带宽未知;
- release/maintenance 压力可能挤压深度研究;
- 项目 ownership、authorship 和维护边界未知;
- public visibility 与 academic rigor 的权重可能因项目而异。
个人风险
- 同时维护过多 Pengyi repo;
- 把架构文档误当实证结果;
- 追求“爆款”先于研究问题和实验可信度;
- 未确认 funding 就做不可逆职业决定。
风险控制
- 一次只运行一个 frozen research question;
- 报告、README 和邮件明确区分 measured / unmeasured;
- 通过小规模合作观察 supervision 与项目机制;
- 正式申请、就业 transition 和研究试验保持可逆并行。
18. 最终建议
结论:值得重点申请,也值得尽快进行一次高质量、证据驱动的研究交流。
但目标不应表述为“加入后做爆款 paper/project”,而应表述为:
用可复现 benchmark 和真实开源贡献,研究 Agent 系统在金融时点、检索、工具 故障与报告事实性上的可靠性边界;如果结论成立,再把它发展成论文和长期系统。
当前状态:
- Target:
P0; - Outreach:
NOT YET SENT; - Formal application:
ACTIVE PREPARATION; - Funding:
UNVERIFIED; - Research results:
UNMEASURED; - Human Gate:
REQUIRED。
19. Primary Sources
PI、团队与机构
- HKU official profile: Chao Huang
- Data Intelligence Lab homepage
- Group members and Join Us
- Official publication list
- HKU IDS Research Seed Funds 2025 result
- HKU IDS newsletter, June 2026
招生
- HKU CDS MPhil/PhD admission
- HKU Graduate School: How to apply
- HKU Graduate School: Supporting documents
- HKU Graduate School: HKPFS
论文
- Understanding DeepResearch via Reports
- AutoAgent
- LightRAG
- MiniRAG
- VideoRAG
- AI-Researcher repository and paper entry
- Justice or Prejudice?
开源项目与贡献证据
- HKUDS GitHub organization
- OpenHarness
- Vibe-Trading
- LightRAG PR #3574
- LightRAG PR #3607
- LightRAG PR #3622
- Vibe-Trading PR #1066
20. Snapshot Changelog
2026-08-12 07:14:34 +08:00: captured GitHub organization aggregate and
repository activity snapshot.
2026-08-12 07:16:37 +08:00: froze PI/group/admission/source snapshot and
generated this PPHD full-diligence report.
- Future refresh trigger: supervisor/team page change, new admission deadline,
material paper/project release, funding update, or direct conversation.