- deduplicator: 从直接数据库查询改为调用 /api/webhook/check-duplicates API - database-ingestor: 使用 JSON 文件传参避免 curl 编码问题 - webhook: 修复 tag name 唯一性约束冲突,更新时删除旧关联重建 - 添加 check-duplicates API 端点用于批量去重检测 - 适配 ProjectTag 显式关联表的数据查询逻辑 - ProjectCard: 修复 getProjectIcon 空值处理 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
4.9 KiB
4.9 KiB
name, description, model, color, tools
| name | description | model | color | tools |
|---|---|---|---|---|
| deduplicator | 对来自所有数据源的原始项目进行跨数据源去重。读取工作区的 raw-projects.json,规范化 URL,执行三级去重检测(GitHub URL、Hugging Face URL、Website URL、Slug 匹配),生成新项目列表和任务队列。使用此 agent 当需要对爬取的项目进行去重时。 示例场景: - 多个数据源爬取完成后需要去重 - 检查新项目是否已存在于数据库 输入参数:{"workspace": ".trending-workspace/..."} | inherit | purple | Read, Write, Bash |
统一去重器 Agent
职责
对来自所有数据源的原始项目进行跨数据源去重:
- 读取原始项目数据
- 规范化 URL
- 三级去重检测
- 生成新项目列表
- 生成分析任务队列
输入参数
{
"workspace": ".trending-workspace/..."
}
执行步骤
Step 1: 读取原始数据
从工作区读取 raw-projects.json。
Step 2: URL 规范化
对不同数据源的 URL 进行规范化处理:
function normalizeUrl(url: string): string {
return url.toLowerCase()
.replace(/\/$/, '') // 移除尾部斜杠
.replace(/^https?:\/\//, '') // 移除协议(用于比较)
}
Step 3: 调用去重检查 API
重要:使用 API 接口而非直接访问数据库。
3.1 构造请求体
根据原始项目数据构造 API 请求:
{
"apiKey": "sk_live_agent_park_webhook_key_2025",
"projects": [
{
"githubUrl": "https://github.com/langchain-ai/langchain",
"huggingfaceUrl": null,
"websiteUrl": "https://python.langchain.com",
"slug": "langchain"
}
]
}
URL 提取规则:
- GitHub 项目:
githubUrl= 项目 URL,slug= generateSlug(name) - Hugging Face 模型:
huggingfaceUrl= 模型 URL,slug= generateSlug(name) - Papers with Code:
websiteUrl= 论文/项目 URL,slug= generateSlug(name)
3.2 调用 API
使用 Bash 执行 curl 请求:
curl -X POST http://localhost:3000/api/webhook/check-duplicates \
-H "Content-Type: application/json" \
-d @check-request.json \
-o check-response.json
3.3 处理响应
API 响应格式:
{
"success": true,
"results": [
{
"githubUrl": "https://github.com/langchain-ai/langchain",
"exists": true,
"matchType": "GITHUB_URL",
"projectId": "cm2x8k9d10001",
"projectName": "LangChain"
}
],
"stats": {
"total": 45,
"exists": 10,
"new": 35,
"breakdown": {
"githubUrl": 5,
"huggingfaceUrl": 2,
"websiteUrl": 1,
"slug": 2
}
}
}
MatchType 说明:
GITHUB_URL: 通过 GitHub URL 匹配HUGGINGFACE_URL: 通过 Hugging Face URL 匹配WEBSITE_URL: 通过官网 URL 匹配SLUG: 通过 slug 匹配NONE: 未匹配,新项目
根据 results[i].exists 判断是否为新项目,仅保留 exists: false 的项目。
Step 4: 生成新项目列表
输出 new-projects.json:
{
"metadata": {
"totalRaw": 45,
"duplicates": 10,
"duplicateBreakdown": {
"githubUrl": 5,
"huggingfaceUrl": 2,
"websiteUrl": 1,
"slug": 2
},
"new": 35
},
"projects": [
{
"source": "github",
"name": "langchain-ai/langchain",
"url": "https://github.com/langchain-ai/langchain",
"description": "Building applications with LLMs through composability",
"metadata": { /* ... */ }
}
]
}
Step 5: 生成分析任务队列
输出 task-queue.json:
{
"metadata": {
"totalTasks": 35,
"createdAt": "2025-01-04T12:10:00Z"
},
"tasks": [
{
"id": 1,
"source": "github",
"name": "langchain-ai/langchain",
"url": "https://github.com/langchain-ai/langchain",
"status": "pending"
},
{
"id": 2,
"source": "huggingface",
"name": "meta-llama/Llama-2-7b",
"url": "https://huggingface.co/meta-llama/Llama-2-7b",
"status": "pending"
}
]
}
Slug 生成规则
function generateSlug(name: string): string {
return name
.toLowerCase()
.replace(/[^a-z0-9\s-]/g, '') // 移除特殊字符
.trim()
.replace(/\s+/g, '-') // 空格转连字符
.substring(0, 100) // 限制长度
}
跨数据源去重示例
同一个项目可能同时出现在:
- GitHub Trending:
https://github.com/openai/whisper - Hugging Face:
https://huggingface.co/openai/whisper-large-v3
通过 GitHub URL 匹配,识别为重复项目,保留一个即可。
错误处理
| 场景 | 处理方式 |
|---|---|
| raw-projects.json 不存在 | 错误提示 "请先运行任务派发器" |
| API 调用失败 | 错误提示并终止,保存中间结果到 check-response.json |
| API 返回 success: false | 错误提示 API 错误详情 |
输出
成功后,返回:
- 新项目数量
- 重复项目数量及原因分布
- 任务队列路径