# 数据新增流程设计文档 > 本文档详细描述了 AI 项目导航站的数据获取、处理和维护的完整流程 ## 目录 - [1. 概述](#1-概述) - [2. 数据源发现](#2-数据源发现) - [3. 深度数据提取](#3-深度数据提取) - [4. 数据标准化处理](#4-数据标准化处理) - [5. 质量控制机制](#5-质量控制机制) - [6. 数据维护策略](#6-数据维护策略) - [7. 技术实现架构](#7-技术实现架构) - [8. 实施路径](#8-实施路径) --- ## 1. 概述 ### 1.1 流程全景 ``` ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ Discovery │ -> │ Fetch │ -> │ Parse │ -> │ Validate │ │ 发现项目 │ │ 深度抓取 │ │ 解析标准化 │ │ 质量验证 │ └─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘ ↓ ↓ ↓ ↓ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ 多源种子 │ │ 完整元数据 │ │ 结构化数据 │ │ 质量评分 │ └─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘ ↓ ┌─────────────┐ │ Store │ │ 入库存储 │ └─────────────┘ ↓ ┌─────────────┐ │ Maintain │ │ 持续维护 │ └─────────────┘ ``` ### 1.2 核心目标 - **完整性**:获取项目的多维度信息(基础、统计、内容、关系、更新、社区) - **准确性**:通过多源交叉验证和质量评分确保数据可靠 - **时效性**:定期更新活跃项目,及时下架失效项目 - **可扩展性**:模块化架构便于添加新数据源 --- ## 2. 数据源发现 ### 2.1 主要数据平台 | 平台 | 数据类型 | API 能力 | 数据量级 | |------|---------|---------|---------| | **GitHub** | 代码项目 | REST API + GraphQL | ★★★★★ | | **Hugging Face** | 模型/数据集/空间 | REST API | ★★★★☆ | | **Papers with Code** | 论文+代码 | Web Scraping | ★★★☆☆ | | **arXiv** | 论文预印本 | REST API | ★★★★☆ | | **Product Hunt** | AI产品 | 无公开API | ★★☆☆☆ | | **AI导航站** | 聚合列表 | Web Scraping | ★★☆☆☆ | ### 2.2 项目发现渠道 #### 2.2.1 趋势榜单 ``` GitHub Trending: https://github.com/trending - 参数:since=daily/weekly/monthly, language={python,typescript,...} - 获取:spike_count、stars、forks、description Hugging Face Trending: https://huggingface.co/api/models - 参数:sort=downloads/likes, trending=true - 获取:modelId、downloads、likes、pipeline_tag ``` #### 2.2.2 社区讨论 - **Reddit**: r/MachineLearning、r/artificial、r/LocalLLaMA 高赞帖 - **Hacker News**: AI相关front page故事 - **Twitter/X**: AI influencer(Andrew Ng、Yann LeCun等)转发 - **Discord/Slack**: AI社区热帖 #### 2.2.3 论文与代码关联 ``` Papers with Code: - Tasks分类:https://paperswithcode.com/tasks - Leaderboards: https://paperswithcode.com/leaderboards - 关联GitHub仓库 ``` #### 2.2.4 聚合站点 - FutureTools、There's An AI For That、AI Valley - 需注意:这些站点数据源自上游,去重更重要 ### 2.3 发现优先级 ``` P0: GitHub Trending + Hugging Face Trending(每日) P1: Papers with Code新入库论文(每周) P2: Reddit/HN高赞讨论(每周) P3: Product Hunt AI产品(每月) P4: AI导航站爬取(按需) ``` --- ## 3. 深度数据提取 ### 3.1 提取维度架构 ```yaml 项目元数据模型: 基础层: - 名称: name, nameEn - 描述: description, descriptionEn (100-500字) - 主页: homepage_url - 许可证: license 统计层: - GitHub: stars, forks, watchers, open_issues - Hugging Face: downloads, likes, discussions - 变化趋势: stars_delta_7d, stars_delta_30d 内容层: - README: 完整Markdown内容 (content/contentEn) - 文档: docs_url, wiki_url - 演示: demo_url, video_url 技术层: - 编程语言: languages (按代码量排序) - 依赖项: dependencies (package.json, requirements.txt) - 框架: frameworks (PyTorch, TensorFlow, LangChain...) - 模型类型: LLM, Diffusion, Computer Vision... 关系层: - 作者: author, author_url - 组织: organization, org_url - 相关项目: similar_projects, forks_from 更新层: - 首次发布: created_at - 最后更新: updated_at - 最后提交: pushed_at - 版本历史: releases, tags 社区层: - 贡献者: contributors_count, top_contributors - Issue活动: open_issues, closed_issues, issue_response_time - 讨论质量: discussions_count, avg_engagement ``` ### 3.2 GitHub 深度提取方案 #### 3.2.1 API 调用策略 ```javascript // 推荐使用 GraphQL 一次性获取,减少请求次数 const query = ` query($owner: String!, $name: String!) { repository(owner: $owner, name: $name) { # 基础信息 name description homepageUrl licenseInfo { key name } url # 统计数据 stargazers { totalCount } forks { totalCount } watchers { totalCount } openIssues: issues(states: OPEN) { totalCount } # 内容 readme: object(expression: "HEAD:README.md") { ... on Blob { text } } # 技术 languages(orderBy: {field: SIZE, direction: DESC}, first: 10) { edges { node { name } size } } # 更新时间 createdAt updatedAt pushedAt # 发布版本 releases(last: 5, orderBy: {field: CREATED_AT, direction: DESC}) { nodes { tagName name publishedAt } } # 贡献者 contributors: mentionableUsers(first: 20) { nodes { login name url } } # Topics (标签) repositoryTopics(first: 20) { nodes { topic { name } } } # 依赖关系 defaultBranchRef { target { ... on Commit { history(first: 1) { nodes { ... on Commit { file(path: "package.json") { ... on Blob { text } } } } } } } } } } `; ``` #### 3.2.2 REST API 补充 ```javascript // 获取 issue 活跃度 const issuesActivity = await fetch( `https://api.github.com/repos/${owner}/${repo}/issues?state=all&per_page=100&sort=comments` ); // 获取 star 历史(需第三方服务如 star-history.com) const starHistory = await fetch( `https://api.star-history.com/svg?repos=${owner}/${repo}&type=Date` ); // 获取社区健康度 const communityProfile = await fetch( `https://api.github.com/repos/${owner}/${repo}/community/profile` ); ``` ### 3.3 Hugging Face 深度提取方案 #### 3.3.1 模型 API ```javascript const model = await fetch(`https://huggingface.co/api/models/${modelId}`); // 返回结构 { modelId: "meta-llama/Llama-2-7b", author: "meta-llama", downloads: 5000000, likes: 12000, lastModified: "2024-01-15T00:00:00.000Z", // 标签体系 tags: ["transformers", "pytorch", "llm", "arxiv:2307.12345"], pipeline_tag: "text-generation", // README 中的 YAML Frontmatter cardData: { license: "llama2", tags: ["llm", "generative"], datasets: ["commoncrawl"], metrics: ["perplexity"], model_index: { "text-generation": [ { name: "Llama-2-7b", model: "?" } ] } } } ``` #### 3.3.2 README 解析 Hugging Face 的 README 通常包含结构化的 YAML 元数据: ```yaml --- license: llama2 tags: - llm - generative - text generation datasets: - commoncrawl - c4 metrics: - perplexity --- # Llama 2 7B [Markdown 内容...] ``` 需要解析并合并这些元数据。 ### 3.4 Papers with Code 提取 ```javascript // 该平台无公开API,需网页抓取 const paperPage = await fetch(`https://paperswithcode.com/paper/${paperSlug}`); // 提取字段 { title: "Attention Is All You Need", titleEn: "Attention Is All You Need", authors: ["Ashish Vaswani", ...], published: "2017-06-12", arxiv_id: "1706.03762", pdf_url: "https://arxiv.org/pdf/1706.03762.pdf", // 代码实现 frameworks: ["PyTorch", "TensorFlow"], implementations: [ { name: "Tensor2Tensor", github: "tensorflow/tensor2tensor", stars: 12000 }, { name: "Harvard NLP", github: "harvardnlp/annotated-transformer", stars: 5000 } ], // 任务与指标 tasks: ["machine-translation", "language-modeling"], benchmarks: ["WMT 2014 En-De", "WMT 2014 En-Fr"], sota_scores: { "BLEU": 28.4 } } ``` --- ## 4. 数据标准化处理 ### 4.1 统一数据结构 所有数据源最终转换为 `ProjectInputSchema` 格式: ```typescript interface ProjectInput { // 基础信息(必填) name: string; // 中文名称(如无则翻译) nameEn: string; // 英文名称 description: string; // 中文描述(100-500字) descriptionEn: string; // 英文描述 slug: string; // URL友好标识符 // 内容(必填) content: string; // 中文README(Markdown) contentEn: string; // 英文README // 状态 status: "ACTIVE" | "ARCHIVED"; source: "GITHUB" | "HUGGING_FACE" | "PAPERS_WITH_CODE" | "MANUAL"; // 关联(必填) tags: string[]; // 1-10个标签 externalLinks: ExternalLink[]; // 1-10个链接 } ``` ### 4.2 标签智能生成 #### 4.2.1 标签分类体系 ```yaml 技术栈标签: - 来源: GitHub languages, HF tags - 示例: Python, TypeScript, PyTorch, TensorFlow 应用领域标签: - 来源: README关键词, HF pipeline_tag, PwC tasks - 示例: Computer Vision, NLP, Reinforcement Learning 模型类型标签: - 来源: README, paper tags - 示例: LLM, Diffusion, GAN, Transformer 框架标签: - 来源: dependencies, README - 示例: LangChain, Gradio, Streamlit, FastAPI 商业状态标签: - 来源: license, homepage - 示例: Open Source, Commercial, Research Only ``` #### 4.2.2 标签提取算法 ```javascript async function extractTags(project) { const tags = new Set(); // 1. 从平台标签直接获取 if (project.githubTopics) { project.githubTopics.forEach(t => tags.add(normalizeTag(t))); } // 2. 从 HF pipeline_tag 获取 if (project.pipelineTag) { tags.add(normalizeTag(project.pipelineTag)); } // 3. NLP 关键词提取(使用NER) const keywords = await extractKeywords(project.descriptionEn); keywords.forEach(kw => { if (isTechnicalTerm(kw)) tags.add(normalizeTag(kw)); }); // 4. 编程语言映射 if (project.languages) { Object.keys(project.languages).forEach(lang => { tags.add(normalizeTag(lang)); }); } // 5. 去重与标准化 return Array.from(tags) .filter(t => t.length >= 2 && t.length <= 30) .map(t => applyTagAlias(t)); // "LLM" -> "Large Language Model" } ``` #### 4.2.3 标签标准化规则 ```javascript const tagAliases = { "LLM": "Large Language Model", "llm": "Large Language Model", "GPT": "Generative Pre-trained Transformer", "CV": "Computer Vision", "NLP": "Natural Language Processing" }; const tagSynonyms = { "diffusion": ["stable-diffusion", "ddpm", "score-based"], "transformer": ["attention", "self-attention"], "fine-tuning": ["finetuning", "fine_tuning"] }; ``` ### 4.3 多语言内容生成 #### 4.3.1 翻译策略 ```javascript async function translateProject(project, sourceLang, targetLang) { // 1. 名称翻译(保留专有名词) const translatedName = await translateText(project.name, { preserveTerms: ["Transformer", "Diffusion", "LLaMA"], format: "title" }); // 2. 描述翻译 const translatedDesc = await translateText(project.description, { maxLength: 500, preserveFormatting: true }); // 3. README 分段翻译 const translatedContent = await translateMarkdown(project.content, { skipCodeBlocks: true, preserveLinks: true, preserveImages: true }); return { name: targetLang === 'zh' ? translatedName : project.name, nameEn: targetLang === 'en' ? translatedName : project.name, // ... }; } ``` #### 4.3.2 翻译质量检查 ```javascript function validateTranslation(original, translated) { const checks = { lengthRatio: translated.length / original.length, // 异常检测:中译英应在0.6-1.5倍之间 hasPreservedTerms: original.match(/[A-Z]{2,}/g).every(term => translated.includes(term) ), noBrokenFormatting: !translated.includes('```') || translated.match(/```/g).length % 2 === 0, noImageLoss: (original.match(/!\[.*\]\(.*\)/g) || []).length === (translated.match(/!\[.*\]\(.*\)/g) || []).length }; return Object.values(checks).every(v => v === true); } ``` ### 4.4 Slug 生成规则 ```javascript function generateSlug(name, nameEn, existingSlugs) { // 1. 优先使用英文 let slug = nameEn .toLowerCase() .replace(/[^a-z0-9]+/g, '-') .replace(/^-|-$/g, ''); // 2. 检查冲突 let finalSlug = slug; let counter = 1; while (existingSlugs.includes(finalSlug)) { finalSlug = `${slug}-${counter}`; counter++; } return finalSlug; } ``` --- ## 5. 质量控制机制 ### 5.1 质量评分模型 ```javascript function calculateQualityScore(project) { const scores = { completeness: 0, // 完整性 (0-30) freshness: 0, // 时效性 (0-25) activity: 0, // 活跃度 (0-25) authority: 0, // 权威性 (0-10) usability: 0 // 可用性 (0-10) }; // 1. 完整性评分 (30分) if (project.name && project.description) scores.completeness += 10; if (project.content && project.content.length > 500) scores.completeness += 10; if (project.externalLinks.length >= 2) scores.completeness += 5; if (project.tags.length >= 3) scores.completeness += 5; // 2. 时效性评分 (25分) const daysSinceUpdate = (Date.now() - new Date(project.updatedAt)) / (1000 * 60 * 60 * 24); if (daysSinceUpdate < 30) scores.freshness = 25; else if (daysSinceUpdate < 90) scores.freshness = 20; else if (daysSinceUpdate < 180) scores.freshness = 15; else if (daysSinceUpdate < 365) scores.freshness = 10; else scores.freshness = 5; // 3. 活跃度评分 (25分) const stars = project.stars || 0; if (stars > 10000) scores.activity += 10; else if (stars > 1000) scores.activity += 7; else if (stars > 100) scores.activity += 5; else if (stars > 10) scores.activity += 3; const recentCommits = project.recentCommits || 0; if (recentCommits > 10) scores.activity += 15; else if (recentCommits > 5) scores.activity += 10; else if (recentCommits > 0) scores.activity += 5; // 4. 权威性评分 (10分) if (project.isOfficialOrg) scores.authority += 5; if (project.hasPaperBacking) scores.authority += 3; if (project.stars > 5000) scores.authority += 2; // 5. 可用性评分 (10分) if (project.hasInstallationGuide) scores.usability += 4; if (project.hasDemo) scores.usability += 3; if (project.hasDocumentation) scores.usability += 3; // 总分 const totalScore = Object.values(scores).reduce((a, b) => a + b, 0); return { totalScore, breakdown: scores, quality: totalScore >= 70 ? 'HIGH' : totalScore >= 40 ? 'MEDIUM' : 'LOW' }; } ``` ### 5.2 垃圾项目检测 ```javascript function detectSpamProject(project) { const signals = []; // 1. 描述异常相似 if (isDescriptionTemplate(project.description)) { signals.push('template_description'); } // 2. Star 增长异常 const starGrowthRate = project.stars / project.daysSinceCreated; if (starGrowthRate > 1000 && project.daysSinceCreated < 7) { signals.push('suspicious_star_growth'); } // 3. 内容过短 if (project.content.length < 100) { signals.push('minimal_content'); } // 4. 缺少基本链接 if (!project.externalLinks.some(l => l.type === 'GITHUB' || l.type === 'WEBSITE')) { signals.push('missing_repository'); } // 5. 关键词堆砌 const keywordDensity = calculateKeywordDensity(project.description); if (keywordDensity > 0.3) { signals.push('keyword_stuffing'); } return { isSpam: signals.length >= 3, signals, confidence: signals.length / 5 }; } ``` ### 5.3 去重策略 利用现有 Webhook 的多级去重机制: ```javascript async function deduplicateProject(newProject) { const { githubUrl, websiteUrl, slug } = newProject; // P0: GitHub URL 精确匹配 const githubMatch = await prisma.externalLink.findUnique({ where: { url_type: { url: githubUrl, type: 'GITHUB' } }, include: { project: true } }); if (githubMatch) { return { exists: true, project: githubMatch.project, reason: 'GITHUB_URL' }; } // P1: Website URL 精确匹配 if (websiteUrl) { const websiteMatch = await prisma.externalLink.findUnique({ where: { url_type: { url: websiteUrl, type: 'WEBSITE' } }, include: { project: true } }); if (websiteMatch) { return { exists: true, project: websiteMatch.project, reason: 'WEBSITE_URL' }; } } // P2: Slug 匹配 const slugMatch = await prisma.project.findUnique({ where: { slug } }); if (slugMatch) { return { exists: true, project: slugMatch, reason: 'SLUG' }; } return { exists: false }; } ``` --- ## 6. 数据维护策略 ### 6.1 增量更新机制 ```javascript // 更新优先级 const UPDATE_PRIORITIES = { HIGH: { interval: '7d', condition: 'stars > 1000 && updated < 7d ago' }, MEDIUM: { interval: '30d', condition: 'stars > 100 && updated < 30d ago' }, LOW: { interval: '90d', condition: 'stars <= 100' } }; async function scheduleUpdate(project) { const priority = determineUpdatePriority(project); // 使用 BullMQ 队列 await updateQueue.add('refresh-project', { projectId: project.id, source: project.source }, { delay: parseInterval(priority.interval), attempts: 3, backoff: { type: 'exponential', delay: 5000 } }); } ``` ### 6.2 生命周期管理 ```javascript async function manageProjectLifecycle(project) { const daysSinceUpdate = (Date.now() - new Date(project.updatedAt)) / (1000 * 60 * 60 * 24); // 1. 活跃项目(90天内更新) if (daysSinceUpdate < 90) { await prisma.project.update({ where: { id: project.id }, data: { status: 'ACTIVE' } }); } // 2. 不活跃项目(90-365天) else if (daysSinceUpdate < 365) { // 检查是否仍在维护 const stillActive = await checkMaintenanceStatus(project); if (!stillActive) { await prisma.project.update({ where: { id: project.id }, data: { status: 'ARCHIVED' } }); } } // 3. 长期未更新(超过365天) else { await prisma.project.update({ where: { id: project.id }, data: { status: 'ARCHIVED' } }); } } ``` ### 6.3 死链检测 ```javascript async function checkExternalLinks() { const links = await prisma.externalLink.findMany(); for (const link of links) { try { const response = await fetch(link.url, { method: 'HEAD', timeout: 5000 }); if (response.status === 404) { // 标记失效 await prisma.externalLink.update({ where: { id: link.id }, data: { valid: false } }); } else if (response.status >= 400) { // 标记异常 await prisma.externalLink.update({ where: { id: link.id }, data: { valid: false, lastError: response.status } }); } } catch (error) { // 网络错误,标记待重检 await prisma.externalLink.update({ where: { id: link.id }, data: { lastCheckFailed: true } }); } } // 移除长期失效的链接 await prisma.externalLink.deleteMany({ where: { valid: false, updatedAt: { lt: new Date(Date.now() - 30 * 24 * 60 * 60 * 1000) } } }); } ``` ### 6.4 热度衰减算法 ```javascript function calculateTrendingScore(project) { const BASE_SCORE = project.stars || 0; // 时间衰减(半衰期30天) const daysSinceUpdate = (Date.now() - new Date(project.updatedAt)) / (1000 * 60 * 60 * 24); const timeDecay = Math.pow(0.5, daysSinceUpdate / 30); // 增长加权(最近7天的star增长) const recentGrowth = (project.stars - project.stars7dAgo) || 0; const growthBonus = recentGrowth * 2; // 社区活跃度 const activityBonus = (project.recentCommits || 0) * 10 + (project.issuesClosedLastWeek || 0) * 5; return (BASE_SCORE * timeDecay) + growthBonus + activityBonus; } ``` --- ## 7. 技术实现架构 ### 7.1 系统架构 ``` ┌─────────────────────────────────────────────────────────────┐ │ 调度层 │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ 定时任务 │ │ 事件触发 │ │ 手动触发 │ │ │ │ (cron) │ │ (webhook) │ │ (admin) │ │ │ └──────────────┘ └──────────────┘ └──────────────┘ │ └─────────────────────────────────────────────────────────────┘ ↓ ┌─────────────────────────────────────────────────────────────┐ │ 采集层 │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ GitHub │ │ Hugging Face │ │ Papers w/ │ │ │ │ Adapter │ │ Adapter │ │ Code Adapter │ │ │ └──────────────┘ └──────────────┘ └──────────────┘ │ └─────────────────────────────────────────────────────────────┘ ↓ ┌─────────────────────────────────────────────────────────────┐ │ 解析层 │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ 数据标准化 │ │ 标签生成 │ │ 多语言翻译 │ │ │ │ (normalizer) │ │ (tagger) │ │ (translator)│ │ │ └──────────────┘ └──────────────┘ └──────────────┘ │ └─────────────────────────────────────────────────────────────┘ ↓ ┌─────────────────────────────────────────────────────────────┐ │ 验证层 │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ 质量评分 │ │ 去重检测 │ │ 垃圾过滤 │ │ │ │ (scorer) │ │ (deduper) │ │ (spam-filter)│ │ │ └──────────────┘ └──────────────┘ └──────────────┘ │ └─────────────────────────────────────────────────────────────┘ ↓ ┌─────────────────────────────────────────────────────────────┐ │ 存储层 │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ │ │ Webhook │ │ Prisma │ │ PostgreSQL │ │ │ │ API │ │ ORM │ │ Database │ │ │ └──────────────┘ └──────────────┘ └──────────────┘ │ └─────────────────────────────────────────────────────────────┘ ↓ ┌─────────────────────────────────────────────────────────────┐ │ 任务队列 │ │ ┌──────────────────────────────┐ │ │ │ BullMQ Queue │ │ │ │ - 采集任务 │ │ │ │ - 更新任务 │ │ │ │ - 死链检测 │ │ │ └──────────────────────────────┘ │ └─────────────────────────────────────────────────────────────┘ ``` ### 7.2 目录结构 ``` src/ ├── lib/ │ ├── scrapers/ # 数据采集器 │ │ ├── base.ts # 基础采集器接口 │ │ ├── github.ts # GitHub采集器 │ │ ├── huggingface.ts # HF采集器 │ │ └── paperswithcode.ts │ │ │ ├── processors/ # 数据处理器 │ │ ├── normalizer.ts # 数据标准化 │ │ ├── tagger.ts # 标签生成 │ │ ├── translator.ts # 多语言翻译 │ │ └── slugify.ts # Slug生成 │ │ │ ├── validators/ # 数据验证器 │ │ ├── scorer.ts # 质量评分 │ │ ├── deduper.ts # 去重检测 │ │ └── spam-filter.ts # 垃圾过滤 │ │ │ └── queue/ # 任务队列 │ ├── producer.ts # 任务生产者 │ ├── consumer.ts # 任务消费者 │ └── jobs/ # 任务定义 │ ├── fetch-project.ts │ ├── refresh-project.ts │ └── check-links.ts │ ├── app/ │ └── api/ │ └── admin/ # 管理API │ ├── ingest/ │ │ └── route.ts # 手动触发采集 │ └── maintenance/ │ └── route.ts # 手动触发维护 │ └── scripts/ ├── ingest-trending.ts # 采集trending项目 ├── refresh-all.ts # 刷新所有项目 └── health-check.ts # 系统健康检查 ``` ### 7.3 核心接口定义 #### 7.3.1 采集器接口 ```typescript // src/lib/scrapers/base.ts export interface ProjectScraper { // 识别平台 platform: ProjectSource; // 从URL识别是否属于该平台 canHandle(url: string): boolean; // 获取项目基础信息 fetchBasic(url: string): Promise; // 获取项目完整信息 fetchFull(url: string): Promise; // 获取趋势列表 fetchTrending(options?: TrendingOptions): Promise; } export interface BasicProjectInfo { name: string; description: string; homepage?: string; repository: string; stars?: number; } export interface FullProjectInfo extends BasicProjectInfo { content: string; languages: Record; tags: string[]; contributors: number; lastUpdated: Date; // ... } ``` #### 7.3.2 处理器接口 ```typescript // src/lib/processors/normalizer.ts export async function normalizeProject( rawProject: FullProjectInfo, source: ProjectSource ): Promise { // 1. 基础字段映射 const base = { name: rawProject.name, nameEn: rawProject.name, description: rawProject.description, descriptionEn: rawProject.description, // ... }; // 2. 内容处理 const content = processMarkdown(rawProject.content); // 3. 标签生成 const tags = await extractTags(rawProject); // 4. Slug生成 const slug = generateSlug(base.name, base.nameEn); // 5. 多语言翻译 const translated = await translateIfNeeded(base, content); return { ...base, ...translated, slug, tags, content, source, externalLinks: buildExternalLinks(rawProject), status: 'ACTIVE' }; } ``` ### 7.4 任务队列配置 ```typescript // src/lib/queue/producer.ts import { Queue } from 'bullmq'; import Redis from 'ioredis'; const connection = new Redis({ host: process.env.REDIS_HOST, port: 6379, maxRetriesPerRequest: 3 }); export const ingestQueue = new Queue('project-ingestion', { connection }); export async function scheduleIngest(url: string) { await ingestQueue.add('ingest-project', { url }, { attempts: 3, backoff: { type: 'exponential', delay: 5000 }, removeOnComplete: { count: 1000 }, removeOnFail: { count: 5000 } }); } export async function scheduleBulkIngest(urls: string[]) { const jobs = urls.map(url => ({ name: 'ingest-project', data: { url } })); await ingestQueue.addBulk(jobs); } ``` ```typescript // src/lib/queue/consumer.ts import { Worker } from 'bullmq'; import { scrapeProject } from '../scrapers'; import { normalizeProject } from '../processors/normalizer'; import { validateProject } from '../validators'; import { prisma } from '../prisma'; const worker = new Worker('project-ingestion', async (job) => { const { url } = job.data; // 1. 识别平台并采集 const scraper = identifyScraper(url); const rawProject = await scraper.fetchFull(url); // 2. 标准化处理 const normalized = await normalizeProject(rawProject, scraper.platform); // 3. 质量验证 const validation = await validateProject(normalized); if (!validation.passed) { throw new Error(`Validation failed: ${validation.reasons.join(', ')}`); } // 4. 去重检测 const existing = await checkDuplicate(normalized); if (existing.exists) { return { action: 'skipped', reason: 'duplicate', projectId: existing.project.id }; } // 5. 写入数据库(通过Webhook API) const response = await fetch(`${process.env.NEXT_PUBLIC_APP_URL}/api/webhook/projects`, { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ apiKey: process.env.WEBHOOK_API_KEY, projects: [normalized] }) }); if (!response.ok) { throw new Error(`Failed to store project: ${response.statusText}`); } return { action: 'created', projectId: result.id }; }, { connection }); ``` --- ## 8. 实施路径 ### 8.1 阶段规划 #### Phase 1: GitHub MVP(Week 1-2) - [ ] 实现 GitHub 采集器(REST + GraphQL) - [ ] 实现基础标签提取(topics + languages) - [ ] 实现数据标准化流程 - [ ] 集成现有 Webhook API - [ ] 添加基础质量评分 **交付物**:能从 GitHub URL 采集完整项目信息 #### Phase 2: Trending 自动化(Week 3) - [ ] 实现 GitHub Trending 解析 - [ ] 配置定时任务(每日凌晨) - [ ] 实现去重逻辑 - [ ] 添加监控告警 **交付物**:每日自动采集 trending 项目 #### Phase 3: Hugging Face 集成(Week 4) - [ ] 实现 HF 采集器 - [ ] 解析 HF YAML 元数据 - [ ] 实现 HF Trending 采集 - [ ] 扩展标签体系(pipeline_tag) **交付物**:支持 HF 模型/数据集 #### Phase 4: 智能化增强(Week 5-6) - [ ] 实现 NLP 标签提取 - [ ] 集成翻译 API(DeepL 或 GPT-4) - [ ] 实现质量评分模型 - [ ] 添加垃圾项目检测 **交付物**:自动化标签生成和翻译 #### Phase 5: 维护系统(Week 7) - [ ] 实现增量更新机制 - [ ] 实现死链检测 - [ ] 实现热度衰减算法 - [ ] 添加生命周期管理 **交付物**:数据自动维护 #### Phase 6: 扩展数据源(Week 8+) - [ ] Papers with Code 集成 - [ ] Reddit/HN 讨论挖掘 - [ ] Product Hunt 集成 - [ ] AI导航站爬取 **交付物**:多源数据融合 ### 8.2 监控指标 ```yaml 采集指标: - 每日新增项目数: target >= 20 - 采集成功率: target >= 95% - API调用次数: 监控配额使用 质量指标: - 高质量项目占比: target >= 70% - 垃圾项目过滤率: target >= 98% - 去重准确率: target >= 99% 维护指标: - 死链检测覆盖率: 100% - 更新及时性: 活跃项目7天内更新 - 数据新鲜度: 90%项目在90天内更新 ``` ### 8.3 技术选型 ```yaml 任务队列: BullMQ (基于Redis) 定时任务: node-cron 爬虫框架: axios + cheerio NLP处理: OpenAI API / Hugging Face Inference API 翻译服务: DeepL API / OpenAI API 监控告警: Sentry + 自定义webhook ``` --- ## 9. 附录 ### 9.1 API 密钥配置 ```bash # .env.local GITHUB_TOKEN=ghp_xxxxx HUGGING_FACE_TOKEN=hf_xxxxx DEEPL_API_KEY=xxxxx OPENAI_API_KEY=sk-xxxxx WEBHOOK_API_KEY=xxxxx REDIS_HOST=localhost ``` ### 9.2 参考资源 - GitHub REST API: https://docs.github.com/en/rest - GitHub GraphQL API: https://docs.github.com/en/graphql - Hugging Face API: https://huggingface.co/docs/huggingface_hub/guides/huggingface_hub_pipelines - Papers with Code: https://paperswithcode.com/docs/api.html - DeepL API: https://www.deepl.com/docs-api --- **文档版本**: v1.0 **最后更新**: 2024年 **维护者**: AI项目导航站团队