Version: 1.6
Status: Living spec (validated against main at e9898b6 on 2026-07-29)
Target: Local dev first; optional Vercel Blob storage
Engineers preparing to discuss a repository need to find entrypoints, understand structure, and identify high-risk areas quickly. RepoAtlas produces an evidence-backed Candidate Brief for an interview walkthrough, bug investigation, planned change, or pull-request discussion without AI-generated claims.
Three supported starts — the bundled sample, a public GitHub repository URL, or a zip upload — produce a Candidate Brief + Repo Analysis with:
- Candidate Brief – Primary evidence-backed walkthrough output: repository summary, prioritized reading path, supported architecture claims, structural hotspots, detected commands, timed explanations, and evidence index (deterministic, no AI).
- Folder Map – Directory tree of the repo.
- Architecture Map – Interactive dependency graph in the runtime UI (ELK layout + pan/zoom).
- Start Here – Prioritized reading list with explanations.
- Danger Zones – Risk-ranked files/modules with breakdown.
- Run and Contribute – Commands extracted from configs and docs.
- Export – Full report as client-generated PDF or PNG in every completed session. Saved reports also support downloadable Markdown with Mermaid graph artifact text.
- A static analysis tool that ingests repositories and produces structured briefs.
- Supports language packs (TS/JS, Python, Java). Depth is uneven: TS/JS is AST-backed via the TypeScript Compiler API; Python and Java are structured heuristics. Do not describe the three packs as equivalent “deep analysis.”
- Works for any repo at a basic level; provides richer signals for supported languages.
- Not a runtime profiler or debugger.
- Not a security vulnerability scanner.
- Not a CI/CD replacement.
- Does not execute or run repository code.
- Does not use LLMs or external AI APIs — all brief text is deterministic.
RepoAtlas does not assert confirmed bugs or vulnerabilities, production readiness, business purpose beyond extracted README/metadata, code correctness, or reliable dynamic runtime behavior. Danger zones reflect structural risk signals (size, coupling, complexity, test proximity, optional churn) — not defect counts.
Bundled sample OR custom input (public GitHub URL or zip upload) → Validation → Analysis (loading) → Report Tabs
- User starts the bundled sample directly or picks a custom input via an accessible tablist:
- Public GitHub URL — pastes a canonical
https://github.com/owner/repoURL with an optional branch/tag ref (default), or - Upload ZIP — selects a
.zipof the repository.
- Public GitHub URL — pastes a canonical
- Client validates custom input client-side (mirroring server rules) and submits
POST /api/analyze(multipart for zip, JSON{ githubUrl, ref? }for GitHub). The sample uses the server-owned fixture with fixed interview focus and inline private handling. - Server saves zip to temp / downloads the public GitHub archive, extracts, runs analyzer.
- UI shows an honest loading state ("Analyzing… (up to 2 min)").
- On success, the API always returns a valid
reportIdplus a persistence status:- Saved:
{ reportId, persisted: true }; the UI fetches the stored report by ID. - Inline fallback:
{ reportId, report, persisted: false }; the UI renders the returned report and does not expose the unsaved ID as a report capability.
- Saved:
- User can export report views client-side (PDF/PNG) from either result. Markdown requires a saved report.
- Sharing uses a 7-day stored token when the report was saved or a 7-day encrypted browser-only URL fragment when it was returned inline.
The legacy JSON zipRef field is not accepted over the network (see §10); internal code and tests call analyzeRepository() directly for that path.
Public product and trust pages:
| Route | Purpose |
|---|---|
/ |
Analysis input, bundled sample, and completed saved or inline report workspace |
/interview-preparation |
Interview-preparation use case with a direct path to analysis |
/code-review-interview, /codebase-interview-preparation, /how-to-walk-through-a-project-in-an-interview, /repository-walkthrough-interview |
SEO/guide variants of the interview-preparation journey, each with a bounded start panel |
/privacy |
Repository and report handling boundaries |
/terms |
Service and output interpretation boundaries |
/contact |
Managed support contact |
/report/:id |
Legacy direct view for a saved report |
/share/:token |
Read-only stored-token or portable encrypted share view |
Completed report workspace at /, /report/:id, or /share/:token:
| Tab | Content | Acceptance Criteria |
|---|---|---|
| Candidate Brief | Timed walkthroughs, repo summary, evidence-backed interview answers, reading path, first PR plan, resume bullets, evidence | Default tab; repository-specific claims link to evidence refs |
| Overview | Repo metadata, deep analysis panels, share link, run commands summary | Shows project_profile, test_inventory, architecture_insights, commit_insights when present; partial badge when timed out |
| Folder Map | Recursive tree with expand/collapse | Renders folder_map; depth limit respected |
| Architecture Map | Interactive ELK-based dependency graph (pan/zoom) | Renders architecture; collapses if nodes > 50; zero-edge reports state the evidence limit and recovery paths |
| Start Here | Sortable table: path, score, explanation | Sorted by score desc; explanations visible |
| Danger Zones | Sortable table: path, score, breakdown | Sorted by score desc; metrics breakdown visible |
| Run & Contribute | Run commands + contribute signals (docs, CI) | Lists commands with source; lists found docs/CI |
| Export | PDF, PNG, Markdown availability, and private sharing | PDF/PNG work for saved and inline reports; Markdown is enabled only for saved reports; share failures offer PDF recovery |
- Idle: Bundled sample action visible; custom ZIP / GitHub URL inputs remain available.
- Analyzing: A single honest indicator ("Analyzing… (up to 2 min)"). No fabricated staged progress — the analyzer does not stream stage events, so the UI does not pretend to. Submit is disabled.
- Fetching saved report: After
{ persisted: true }and a validatedreportIdare returned, an optional skeleton for tabs until the stored report loads. - Rendering inline report: After
{ persisted: false, report }is returned, the UI renders the report immediately without a second API request.
| Error | User-facing message | HTTP/Code |
|---|---|---|
| Invalid input | "Provide a GitHub repository URL, upload a zip file, or request the sample." | 400 / INVALID_INPUT |
| Invalid URL | "Enter a canonical GitHub repository URL like https://github.com/owner/repo." | 400 / INVALID_URL |
| Repository not found | "Repository not found (it may be private)." | 404 / REPO_NOT_FOUND |
| Private repository | "Repository is private." | 403 / REPO_PRIVATE |
| Missing ref | "Requested branch or tag was not found." | 404 / MISSING_REF |
| GitHub rate limit | "GitHub rate limit reached. Try again later." | 429 / RATE_LIMITED |
| Download timeout | "GitHub request timed out." | 504 / DOWNLOAD_TIMEOUT |
| Repo too large | "Repository archive exceeds the size limit." | 413 / REPO_TOO_LARGE |
| Zip/archive invalid | "Invalid or corrupted zip file." | 400 / ZIP_INVALID |
| Timeout | "Analysis timed out. Try a smaller repository." | 504 / TIMEOUT — or partial report saved with partial: true when indexing completed before deadline |
| Too many requests | "Too many analysis requests. Please wait and try again." | 429 / RATE_LIMIT_EXCEEDED |
| Analysis failed | "Analysis failed. Check server logs." | 500 / ANALYSIS_FAILED |
| Zip path not found | "Zip path not found. Check the path or re-upload." | 404 / ZIP_NOT_FOUND (internal path only) |
- UI supports client-side export workflows (PDF/PNG full-report raster snapshots via
html2canvas+jspdf) for saved and inline reports. - Markdown export is available at
GET /api/reports/:id/export/mdonly after successful persistence. The UI disables it for inline reports and explains the storage requirement. - Saved-report sharing:
POST /api/reports/:id/share→/share/:token(7-day TTL; report JSON only). - Inline-report sharing: the browser compresses and AES-GCM-encrypts the validated report into
/share/portable#…. The URL fragment is not sent to the server; the recipient browser decrypts it and enforces the encoded 7-day expiry. URLs over 24,000 characters are rejected with PDF fallback guidance.
flowchart TB
subgraph UI [Next.js UI]
Input[Input Form]
Tabs[Report Tabs]
Export[PDF + PNG Export]
PortableShare[Encrypted URL Fragment]
end
subgraph API [API Routes]
Analyze[POST /api/analyze]
end
subgraph Worker [Analyzer Worker]
Ingest[Repo Ingest]
Index[Indexing Pipeline]
Packs[Language Packs]
Scoring[Start Here + Danger Zones]
end
subgraph Storage [Storage]
TempWorkspace[Temp Workspace]
ReportJSON[(Filesystem or private Blob JSON)]
end
Input --> Analyze
Analyze --> Ingest
Ingest --> TempWorkspace
Ingest --> Index
Index --> Packs
Packs --> Scoring
Scoring --> Analyze
Analyze -. best-effort save .-> ReportJSON
ReportJSON --> Tabs
Analyze -->|inline fallback| Tabs
Tabs --> Export
Tabs --> PortableShare
| Component | Description |
|---|---|
| Next.js UI | React + TypeScript + Tailwind CSS. Public product and trust pages plus tabbed report views. |
| API Routes | Next.js Route Handlers for analysis, saved reports, Markdown export, stored-token sharing, and cleanup. |
| Analyzer Worker | Node.js analysis hosted in an isolated worker_threads worker by default. Recognized startup failures before the ready handshake may fall back in-process; post-start failures do not rerun the repository. |
| Temp Workspace | os.tmpdir() subdir per analysis. Clone or extract zip here. |
| Report Storage | JSON files on disk locally ({REPORTS_DIR}/{reportId}.json) or in a connected private Vercel Blob store. Production without usable Blob credentials returns reports inline instead of failing the completed analysis. |
| Portable Sharing | Browser-only compression, AES-GCM encryption, expiry validation, and report validation for inline report links. |
- Request: Client
POST /api/analyzewith multipart zip (fileorzip), JSON{ githubUrl, ref? }, or{ sample: true }. JSONzipRefis rejected. - Ingest: Server extracts uploaded zip (server-created temp path only) or downloads a public GitHub archive into temp workspace.
- Analysis: The isolated analyzer worker walks the workspace, runs the common pipeline and applicable language packs, and stops promptly on request abort or deadline.
- Report: Analyzer produces validated
ReportJSON and attempts persistence when the runtime has storage credentials. - Response: Successful persistence returns
{ reportId, persisted: true }. Unavailable storage or a best-effort save failure returns{ reportId, report, persisted: false }. - Render: The UI fetches a saved report by ID or renders the inline report directly. Stored JSON and portable shared JSON are validated before use (
src/lib/reportSchema.ts).
Discriminated input model (src/lib/ingest.ts):
{ kind: "zip", zipRef, zipName? }— server-created temp path (multipart upload) or a server-owned fixture path (sample flow). Never a caller-supplied network path.{ kind: "github", githubUrl, ref? }— a canonical public GitHub URL with an optional validated branch/tag ref.
Accepted URLs — canonical repository URLs only:
https://github.com/owner/repositoryhttps://github.com/owner/repository.git
Rejected: non-HTTPS, non-github.com hosts, tree/blob subpaths, query strings, fragments, non-default ports, and malformed owner/repo. A custom branch/tag is provided through a separate validated ref field (isValidGitRef), not by parsing tree/blob URLs.
Security properties (all enforced and tested):
- Public-only, unauthenticated. GitHub API and archive requests are always sent without an
Authorizationheader. The serverGITHUB_TOKENis never attached to user-supplied URL ingestion, so an unauthenticated caller can never borrow privileged server access to read a private repository. A repository the API reports asprivate: true, or that returns 404/403, is refused before any download (Phase 1 finding B). - Exact-SHA first. The requested ref (or default branch) is resolved to an exact commit SHA via the commits API before downloading. The archive for that exact SHA is downloaded and the same SHA is recorded as
clone_hash(finding C). - Streaming with caps. The archive is streamed to a temp file and aborted if it exceeds
MAX_COMPRESSED_BYTES; it is never buffered whole in memory (finding D). Acontent-lengthover the cap is rejected up front. - Redirect policy. Redirects are followed only when the finally-resolved host is a known GitHub host (
github.com,codeload.github.com,objects.githubusercontent.com); any other host is rejected. - Timeouts. GitHub API requests use
GITHUB_API_TIMEOUT_MS; archive downloads useDOWNLOAD_TIMEOUT_MS. Aborts map toDOWNLOAD_TIMEOUT. - Cleanup. The per-analysis temp directory is removed on success, failure, and cancellation.
Error taxonomy (finding E): INVALID_URL, REPO_NOT_FOUND, REPO_PRIVATE, MISSING_REF, RATE_LIMITED, DOWNLOAD_TIMEOUT, REPO_TOO_LARGE, ZIP_INVALID.
Acceptance criteria: https://github.com/vercel/next.js accepted; https://github.com/vercel/next.js/tree/canary, https://gitlab.com/foo/bar rejected. See src/lib/github.test.ts and src/lib/ingest.github.test.ts (mocked API/archive — CI never hits live GitHub).
- Endpoint: Multipart form upload to
POST /api/analyzewithfileorzipfield. - Compressed limit (environment-aware):
- Vercel deployment:
MAX_DEPLOYED_ZIP_BYTES= 4 MB (Function body cap ~4.5 MB). Direct users to GitHub URL mode for larger public repos. - Local dev:
MAX_COMPRESSED_BYTES= 100 MB. - Selected by
maxCompressedBytesForZipUpload()insrc/lib/ingestLimits.ts.
- Vercel deployment:
- Validation: Check magic bytes
50 4B 03 04or50 4B 05 06(PK) for zip. - Extraction: production ingest streams from disk through
yauzlinsrc/lib/safeZipExtract.ts; the small buffer-based test helper retainsadm-zip. Both paths enforce the path jail, normalized-target collision preflight, entry count cap, per-file cap, and uncompressed size cap. - Path traversal: For each entry, resolve path relative to extract root; reject if resolved path is outside root or contains
... - Path collisions: Reject duplicate normalized destinations (including dot-segment and platform separator aliases) and file/child-path conflicts before writing any entry.
- Size limit: Max cumulative uncompressed 50 MB; abort extraction if exceeded.
Acceptance criteria: Valid zip extracts; zip with ../../../etc/passwd entries, normalized duplicate destinations, or file/child-path conflicts is rejected; oversized zip is aborted. See src/lib/safeZipExtract.test.ts.
- Delete temp dir on analysis completion (success or failure).
analyzeRepository()wraps the whole run intry/finallyand always invokesworkspace.cleanup(), including when report persistence throws (regression test:src/analyzer/cleanup.test.ts). - Report files: filesystem and Vercel Blob storage support deletion and a TTL/max-count sweep.
sweepExpiredReports()usesanalyzed_aton the filesystem and blob upload timestamps for blob storage. Retention isREPORT_TTL_DAYS(default 7 with Blob storage credentials, otherwise 30) andREPORT_MAX_COUNT(default 100). - Share tokens: 7-day TTL;
sweepExpiredShareTokens()lists and deletes expired records on both filesystem and Blob (src/lib/sharing.ts). - Cron:
GETandPOST /api/cron/cleanuprun report + share sweeps. Vercel schedulesGETdaily at 03:00 UTC. In production (NODE_ENV=productionorVERCEL=1), both methods fail closed with503 MISCONFIGUREDwhenCRON_SECRETis unset; when set, they requireAuthorization: Bearer <CRON_SECRET>. - Production without usable Blob credentials returns reports inline. Saved report retention and stored-token cleanup cannot operate until both Blob storage and the protected cleanup schedule are connected.
All size/count/timeout budgets live in one module so the API route, ZIP extractor, GitHub downloader, indexing pipeline, and UI cannot disagree (finding D / Phase 2 requirement 8).
| Limit | Constant | Value | Behavior |
|---|---|---|---|
| Max ZIP upload (Vercel) | MAX_DEPLOYED_ZIP_BYTES |
4 MB | Multipart body cap on deployed Functions |
| Max compressed archive (local ZIP + GitHub) | MAX_COMPRESSED_BYTES |
100 MB | Abort upload/download if exceeded |
| Max uncompressed total | MAX_UNCOMPRESSED_BYTES |
50 MB | Abort extraction if exceeded |
| Max entries | MAX_ENTRIES |
10,000 | Abort extraction |
| Max single file | MAX_SINGLE_FILE_BYTES |
10 MB | Abort extraction |
| Max indexed files | MAX_FILE_COUNT |
10,000 | Stop indexing; add warning |
| Max folder depth | MAX_DEPTH |
10 | Stop recursing; mark node truncated + add warning (no longer silent) |
| Max analysis time | MAX_ANALYSIS_TIME_MS |
120 s | Abort; return partial report |
| Archive download timeout | DOWNLOAD_TIMEOUT_MS |
60 s | Abort → DOWNLOAD_TIMEOUT |
| GitHub API timeout | GITHUB_API_TIMEOUT_MS |
15 s | Abort → DOWNLOAD_TIMEOUT |
| Step | Description | Output |
|---|---|---|
| Folder tree | Recursive fs.readdirSync with depth limit (MAX_DEPTH); over-depth directories are marked truncated and a warning is emitted |
FolderMapNode |
| File metadata | For each file: path, size, extension | FileMetadata[] |
| Language detection | Extension → language map; .gitattributes overrides if present |
language per file |
| Key docs discovery | README*, CONTRIBUTING*, LICENSE*, CHANGELOG*, deterministically sorted (root docs first, then lexicographic) |
keyDocs: string[] |
| CI discovery | Glob: .github/workflows/*.yml, .gitlab-ci.yml, Jenkinsfile |
ciConfigs: string[] |
| Run command extraction | Parse package.json scripts, Makefile, pyproject.toml, pom.xml, build.gradle, Docker Compose, README fenced blocks |
RunCommand[] via src/analyzer/commands/index.ts |
To stop duplicated documentation from producing repetitive or misleading briefs (Phase 3), a deterministic discovery step builds a document_inventory:
- Classify & prioritize. Each doc is classified (
readme,contributing,architecture,docs,changelog,license) and scoped (root,docs,nested). Root docs outrankdocs/, which outrank nested package docs. - Deterministic order. Documents are sorted by (category, scope, path depth, path) before any grouping, independent of filesystem traversal order.
- Duplicate detection. Content is hashed both raw (
content_hash) and normalized (normalized_hash). Normalization strips a UTF-8 BOM, converts CRLF/CR → LF, trims trailing whitespace per line, and collapses surrounding blank lines. Documents sharing a normalized hash form a duplicate group; the highest-priority member iscanonical, others recordduplicate_of. - Similarity flagging. Same-category canonical docs with Jaccard line similarity ≥ 0.85 (and < 1) are recorded in
similar_groupsas possible redundancy — never suppressed. - Nothing hidden. Every document remains in the inventory and folder map. Duplicates are grouped, not deleted, so users can see what was suppressed and why.
- Canonical selection. One canonical document per duplicate group is used for purpose extraction, run-command extraction, and Candidate Brief evidence, so equivalent content does not generate repeated evidence cards. Different nested package READMEs are treated as legitimate and each keep their own card.
Purpose extraction (src/analyzer/purpose.ts) prefers the canonical README, and will not use a heading that is only the repository name; it falls back to the first meaningful paragraph or a manifest description, preserving the source path and evidence.
Tests: src/analyzer/docs.test.ts, src/analyzer/docs.integration.test.ts, src/analyzer/purpose.test.ts, src/analyzer/pipeline.test.ts. Fixture: fixtures/repo-docs-dedup.
| Aspect | Rules |
|---|---|
| Import extraction | TypeScript Compiler API AST walk (src/analyzer/packs/tsjsExtract.ts). Collects static import, side-effect imports, import(), require(), export … from, export * from, and type-only import/export. Comments and string literals never produce edges. |
| Module resolution | ts.resolveModuleName with tsconfig.json / jsconfig.json (baseUrl, paths), relative/extensionless/index resolution, package.json main/module/exports, and npm/pnpm workspace package names (src/analyzer/packs/tsjsResolve.ts). Never leaves the extracted workspace. |
| Semantic graph | Language-neutral semantic_graph (optional on Report, report_version 3+) stores nodes/edges with resolution status, line-bounded evidence, and stats. Folder-level architecture is derived from resolved_internal edges only. Unresolved edges are recorded and must not inflate fan-in/fan-out. External package edges are recorded separately. |
| Entrypoint heuristics | Next.js App Router page/layout/route/middleware, package main/module/browser/bin/exports, narrow dev/start/build script path literals, and common `src/index |
| Test proximity | Test files: *.test.{js,ts}, *.spec.{js,ts}, __tests__/*, *.test.{jsx,tsx}. Proximity = same dir or nearest test dir distance. |
| Structural complexity | AST decision-point count (if/for/while/do/switch cases/catch/conditionals/&&/||/??) plus block nesting depth and LOC. Score = branches*3 + nesting*2 + round(loc/40). This is a structural complexity score, not claimed as cyclomatic complexity. |
| Graph collapse | Module = file; folder = directory. Collapse: group nodes by parent dir; edge A→B becomes dir(A)→dir(B). Caps: 50 nodes / 200 edges. |
Acceptance criteria: TS repo with src/index.ts importing ./utils produces a resolved_internal semantic edge and folder architecture edge; fake import text in comments/strings does not; unresolved imports appear in semantic_graph.stats.unresolved without inflating fan-out.
| Aspect | Rules |
|---|---|
| Import extraction | Import-statement scanner that skips strings and comments, handles parenthesized and line-continued statements, expands from pkg import name, and resolves absolute or relative package paths without claiming a full Python AST. |
| Entrypoint heuristics | if __name__ == "__main__"; setup.py entry_points; pyproject.toml [project.scripts]; -m targets from docs. |
| Test proximity | test_*.py, *_test.py, tests/ dir. |
| Complexity proxy | Line-based decision points aligned with the TS pack definition (if, elif, for, while, except branches, and/or, match/case) + indentation nesting + LOC. else, try, finally, and with are not counted, so equivalent control flow scores comparably across the TS and Python packs (see complexityParity.test.ts). Still a structural proxy, not McCabe cyclomatic complexity. |
| Graph collapse | Module = file; package = dir with __init__.py. |
Acceptance criteria: main.py with from utils import foo produces edge; main.py with if __name__ == "__main__" marked as entrypoint.
| Aspect | Rules |
|---|---|
| Import extraction | Anchored import scanning for ordinary, wildcard, and static imports, with package/class resolution. Safe same-package type references are also recovered after comments and strings are removed. |
| Entrypoint heuristics | public static void main; @SpringBootApplication; JAR manifest Main-Class. |
| Test proximity | *Test.java, *IT.java, src/test/java layout. |
| Complexity proxy | Line-based decision points aligned with the TS pack definition (if, case labels, loops, catch, &&, ||, ternary ?) + brace nesting + LOC. else, the switch keyword itself, and generic wildcards (<?>, <? extends>) are not counted. Still a structural proxy, not McCabe cyclomatic complexity. |
| Graph collapse | Class = file; package = folder. |
Acceptance criteria: Java file with public static void main marked as entrypoint; imports produce edges.
Candidates: Key docs (README first), entrypoint files, code files near entrypoints, build definitions (pom.xml, build.gradle), root configs.
Raw score composition (implemented in src/analyzer/scoring/startHere.ts; weights evolve with regression tests):
-
Root README ≈ 95; other READMEs ≈ 80; CONTRIBUTING ≈ 75; other key docs ≈ 45.
-
Next.js route handlers / app-router entries 65–85; runnable Python entry names (
__main__.py,manage.py,main.py, …) 70–90. -
Fan-in bonus capped at 35; import-distance bonus decays from 90 (distance 0) to single digits by distance 3.
-
Test files receive a −40 penalty.
-
Raw scores are min–max normalized to 0–100 inside
rankStartHere(ties normalize to 50); top 12 candidates are kept, ties broken by path. -
Explanation strings list the contributing signals verbatim.
Metrics:
| Metric | Definition | Source |
|---|---|---|
| Size | LOC or file size in bytes | File metadata |
| Fan-in | Number of files importing this file | Import graph |
| Fan-out | Number of files this file imports | Import graph |
| Complexity | Complexity proxy value | Language pack |
| Test proximity penalty | 0 if nearby test else 1 | Test proximity |
| Churn (optional) | Commit count in last N commits reachable from the ingested tip | Local git log when .git exists; otherwise GitHub commits API scoped with `sha=<clone_hash |
Normalization: For each metric, compute percentile rank (0–100) within repo.
RiskScore formula (test files excluded from ranking; when churn unavailable — default for zip uploads without .git):
RiskScore = (
0.20 * size_percentile +
0.25 * fan_in_percentile +
0.20 * fan_out_percentile +
0.25 * complexity_percentile +
0.10 * (100 - test_proximity_percentile)
)
When commit_insights.mode is local_git or github_api and churn data exists:
RiskScore = (
0.18 * size_percentile +
0.22 * fan_in_percentile +
0.18 * fan_out_percentile +
0.22 * complexity_percentile +
0.10 * (100 - test_proximity_percentile) +
0.10 * churn_percentile
)
See src/analyzer/scoring.ts and adr/003-scoring-semantics.md.
- Clamp to 0–100.
- Sort by RiskScore descending.
Explanation breakdown: e.g. "High fan-in (15), high complexity (42), no nearby tests".
Acceptance criteria: File with high fan-in, high complexity, no tests ranks in top 5 danger zones with correct breakdown.
{
"report_version": 3,
"partial": false,
"repo_metadata": {
"name": "string",
"url": "string",
"branch": "string",
"clone_hash": "string | null",
"analyzed_at": "string (ISO 8601)"
},
"folder_map": { },
"architecture": {
"nodes": [],
"edges": []
},
"semantic_graph": {
"version": 1,
"language": "typescript",
"adapter": "tsjs-typescript-compiler-api",
"nodes": [],
"edges": [],
"stats": {},
"warnings": []
},
"start_here": [],
"danger_zones": [],
"run_commands": [],
"contribute_signals": {
"key_docs": [],
"ci_configs": []
},
"candidate_brief": { },
"project_profile": { },
"project_purpose": { },
"document_inventory": {
"documents": [],
"duplicate_groups": [],
"similar_groups": [],
"canonical_readme": "string | undefined"
},
"technical_decisions": [],
"symbols": [],
"test_inventory": { },
"architecture_insights": { },
"commit_insights": {
"mode": "local_git | github_api | unavailable"
},
"warnings": []
}candidate_brief is the primary interview-facing output. partial: true indicates a timeout after folder map was saved. Deep-analysis fields (project_profile, test_inventory, architecture_insights, commit_insights, semantic_graph) are optional and populated when signals are available. report_version is 3 when semantic_graph may be present; older stored reports without it remain readable.
Stored report JSON is validated at read time in getReport():
- corrupt — missing required fields, wrong types, invalid JSON → treated as not found
- incompatible —
report_versiongreater thanREPORT_VERSION→ treated as not found
New reports are stamped report_version: 3. Older v2 reports without semantic_graph remain readable. A broader migration layer is future work (see roadmap.md).
export interface RepoMetadata {
name: string;
url: string;
branch: string;
clone_hash: string | null;
analyzed_at: string; // ISO 8601
}
export type FolderMapNode = {
path: string;
type: 'file' | 'dir';
children?: FolderMapNode[];
truncated?: boolean; // true when depth-limited entries were not walked
};
export interface ArchitectureNode {
id: string; // file path or module id
label: string; // display name
type?: 'file' | 'module' | 'folder';
}
export interface ArchitectureEdge {
from: string; // node id
to: string; // node id
type?: 'import' | 'dependency';
}
export interface Architecture {
nodes: ArchitectureNode[];
edges: ArchitectureEdge[];
}
export interface StartHereItem {
path: string;
score: number;
explanation: string;
}
export interface DangerZoneItem {
path: string;
score: number;
breakdown: string;
metrics: {
size?: number;
fan_in?: number;
fan_out?: number;
complexity?: number;
test_proximity?: number;
churn?: number;
};
}
export interface RunCommand {
source: string; // e.g. "package.json", "README"
command: string;
description?: string;
}
export interface ContributeSignals {
key_docs: string[];
ci_configs: string[];
}
export interface Report {
report_version?: number;
partial?: boolean;
repo_metadata: RepoMetadata;
folder_map: FolderMapNode;
architecture: Architecture;
semantic_graph?: SemanticGraph;
start_here: StartHereItem[];
danger_zones: DangerZoneItem[];
run_commands: RunCommand[];
contribute_signals: ContributeSignals;
candidate_brief?: CandidateBrief;
project_profile?: ProjectProfile;
project_purpose?: ProjectPurpose;
document_inventory?: DocumentInventory;
technical_decisions?: TechnicalDecision[];
symbols?: CodeSymbol[];
test_inventory?: TestInventory;
architecture_insights?: ArchitectureInsights;
commit_insights?: CommitInsights;
warnings: string[];
}SemanticGraph is defined in src/types/semanticGraph.ts (versioned nodes/edges with resolution status and line-bounded evidence). TS/JS, Python, and Java adapters are implemented and covered by fixtures.
document_inventory (see src/types/report.ts) captures every discovered document, duplicate groups (with canonical + suppressed paths and a reason), optional similar groups, and the chosen canonical_readme.
See src/types/report.ts for full CandidateBrief, EvidenceRef, and deep-analysis type definitions.
| Route file | Methods | Public endpoint | Notes |
|---|---|---|---|
src/app/api/analyze/route.ts |
POST |
/api/analyze |
Accepts multipart upload (file or zip), JSON { "githubUrl", "ref"? }, or JSON { "sample": true }. JSON zipRef is rejected. Rate-limited + concurrency-gated. |
src/app/api/reports/[id]/route.ts |
GET |
/api/reports/:id |
Returns persisted report JSON with Cache-Control: no-store. UUID validated before storage access. No public DELETE — see adr/001-capability-access.md. Stored JSON validated at read time (parseAndValidateReport). |
src/app/api/reports/[id]/share/route.ts |
POST |
/api/reports/:id/share |
Creates 7-day read-only share token |
src/app/api/share/[token]/route.ts |
GET |
/api/share/:token |
Resolves share token to report JSON |
src/app/api/reports/[id]/export/md/route.ts |
GET |
/api/reports/:id/export/md |
Returns text/markdown with attachment headers |
src/app/api/cron/cleanup/route.ts |
GET, POST |
/api/cron/cleanup |
Both methods sweep expired reports + share tokens; Vercel schedules GET daily. Fails closed in production without CRON_SECRET. |
Maintenance rule: when route handlers are added/removed/renamed, update this table in the same PR by checking the route files directly.
Request (zip): multipart/form-data with a single zip file (field file or zip). Max 4 MB on Vercel, 100 MB locally (maxCompressedBytesForZipUpload()).
Request (GitHub): Content-Type: application/json:
{ "githubUrl": "https://github.com/owner/repo", "ref": "optional-branch-or-tag" }Request (sample): Content-Type: application/json with { "sample": true } (analyzes the bundled fixtures/repo-ts).
The old JSON
zipReffield is intentionally rejected with400 / INVALID_INPUT— caller-controlled server paths are never analyzable through the public API (Phase 1 finding A). Internal callers/tests useanalyzeRepository({ zipRef })directly.
Response (200, saved report):
{
"reportId": "uuid-string",
"persisted": true
}Response (200, inline fallback):
{
"reportId": "uuid-string",
"persisted": false,
"report": {
"report_version": 3
}
}The inline response contains the complete validated Report. It is used when persistence is not configured or a best-effort write fails.
Error responses:
| Status | Body | Code |
|---|---|---|
| 400 | { "code": "INVALID_INPUT", "message": "..." } |
Missing/invalid body, unsupported content type, or zipRef sent over the network |
| 400 | { "code": "INVALID_URL", "message": "..." } |
Non-canonical GitHub URL or invalid ref |
| 400 | { "code": "ZIP_INVALID", "message": "..." } |
Invalid zip / archive |
| 403 | { "code": "REPO_PRIVATE", "message": "..." } |
Repository is private/restricted |
| 404 | { "code": "REPO_NOT_FOUND", "message": "..." } |
Repository not found (or private, unauthenticated) |
| 404 | { "code": "MISSING_REF", "message": "..." } |
Requested branch/tag not found |
| 413 | { "code": "REPO_TOO_LARGE", "message": "..." } |
Upload/archive exceeds compressed limit (4 MB on Vercel, 100 MB otherwise) |
| 429 | { "code": "RATE_LIMITED", "message": "..." } |
GitHub API rate limit |
| 429 | { "code": "RATE_LIMIT_EXCEEDED", "message": "..." } |
Per-instance request limit / concurrency cap |
| 504 | { "code": "DOWNLOAD_TIMEOUT", "message": "..." } |
GitHub request/download timed out |
| 504 | { "code": "TIMEOUT", "message": "..." } |
Analysis timeout (may instead return partial: true) |
| 500 | { "code": "ANALYSIS_FAILED", "message": "..." } |
Analysis error |
Returns persisted report JSON. The id must match a strict UUID shape. Loaded JSON is validated via parseAndValidateReport(); corrupt or future-incompatible (report_version > supported) payloads are treated as not found (404), not served as partial data.
There is no public DELETE. Report ids are read-only capabilities; retention is server-side TTL sweep only (adr/001-capability-access.md).
- Current API routes:
POST /api/analyze,GET /api/reports/:id,POST /api/reports/:id/share,GET /api/share/:token,GET /api/reports/:id/export/md, andGET/POST /api/cron/cleanup. There is no publicDELETEroute (removed; retention via server-side TTL sweep). - There is no separate
POST /api/upload; uploads are handled byPOST /api/analyzevia multipart fieldsfileorzip. - Saved reports use
/api/reports/:id,/api/reports/:id/export/md, and the stored-token share routes. Inline reports render from the analysis response, retain PDF/PNG export, and use/share/portable#…for encrypted browser-only sharing.
- Client: No automatic retry. A failed request returns to an actionable input state.
- Server: No automatic retry for clone; single attempt.
| Component | Responsibility |
|---|---|
HomePage |
Coordinates homepage sections, analysis state, sample actions, and completed report focus |
HomepageProofSections |
Outcome-first hero, Candidate Brief outcomes, supported workflows, bundled proof, and trust/FAQ content |
InputForm |
Bundled sample, public GitHub URL, and ZIP submission; calls POST /api/analyze |
AnalysisIntentSelector |
Selects the supported analysis focus without creating separate products |
RepositoryInputControls |
Accessible GitHub/ZIP input tabs and client validation |
ReportTabs |
Tab bar + tab content; receives Report |
CandidateBriefPanel |
Candidate Brief sections with evidence navigation |
DeepAnalysisSection |
Overview panels for project profile, tests, boundaries, commits |
FolderMapTree |
Recursive tree; expand/collapse |
ElkArchitectureGraph |
Interactive ELK graph rendering; collapse to folder if nodes > 50 |
StartHereTable |
Sortable table; path, score, explanation |
DangerZonesTable |
Sortable table; path, score, breakdown |
RunContributeSection |
Lists run commands + contribute signals |
ReportTabs sharing |
Creates a stored-token link for saved reports or an encrypted portable link for inline reports |
- Use
elkjsfor layout and the UI graph component for runtime rendering. - Input:
architecture.nodesandarchitecture.edges. - Generate positioned nodes/edges for an interactive graph view (zoom/pan, fit-to-view).
- Reduction: If
nodes.length > 50, collapse to folder level: group by parent dir; edges between folders. - Fallback: If layout/rendering fails, show raw node/edge list.
- Mermaid is not the runtime UI renderer.
- Mermaid output is retained for markdown artifact rendering/export compatibility.
- Input for markdown export Mermaid remains
architecture.nodesandarchitecture.edges.
if (nodes.length <= 50) use file-level graph
else {
group nodes by directory (e.g. src/utils, src/api)
create folder nodes
edge (A, B) => edge (dir(A), dir(B))
deduplicate edges
}
- Store the saved or inline report in React state for the active page.
- RepoAtlas does not write inline reports to local storage or another persistent browser cache.
- Cross-session access requires a saved report route or an encrypted portable share URL.
| Concern | Mitigation |
|---|---|
| Capability-link access | Report UUID is a read-only capability; public DELETE removed. See adr/001-capability-access.md. |
| Portable share exposure | Inline reports are compressed and AES-GCM-encrypted in the URL fragment. Fragments are not sent in HTTP requests; the recipient browser validates the report and encoded 7-day expiry before rendering. Anyone with the complete link can read it. |
| Browser content execution | Production responses use the tested CSP in securityHeaders.js: same-origin scripts/connections, object-src 'none', frame-src 'none', and no unsafe-eval; data:/blob: are limited to client export image/font/worker capabilities. |
| Arbitrary server file read | Public API rejects JSON zipRef; only server-created temp paths and server-owned fixtures reach the zip ingest path. Report ids validated as UUID before storage access. |
| Corrupt stored JSON | parseAndValidateReport() at read time; invalid payloads return not found. |
| Path traversal (zip) | src/lib/safeZipExtract.ts — resolve paths; reject ..; jail to extract root |
| Code execution | Never require(), import(), or exec() repo code; parse as text only |
| Private repo exposure | User-supplied GitHub ingestion is always unauthenticated — no GITHUB_TOKEN on user requests. Repos reported private, or 404/403, are refused before download. |
| SSRF / redirect abuse | Only canonical github.com URLs are accepted; archive redirects are followed only to known GitHub hosts. |
| Zip bombs / oversized archives | Deployment-aware compressed caps (adr/002-zip-limits.md), uncompressed caps, entry-count and per-file caps, streamed GitHub download with abort-on-cap. |
| Fetch hangs | GitHub API and archive requests have AbortController timeouts. |
| Rate limiting | Process-local concurrency gate (MAX_CONCURRENT_ANALYSES) plus a 30-per-minute sliding window by default. Configured Upstash REST credentials select the distributed limiter once per process; missing or invalid shared-store results use an explicit best-effort path. |
| Cron misconfiguration | Production cleanup returns 503 when CRON_SECRET is unset; both scheduled GET and manual POST require the bearer secret when configured. |
| Share token cleanup on Blob | listShareTokens() and deleteShareRecord() use Vercel Blob list/del with shares/ prefix. |
Security assumptions & remaining limitations:
- Only public GitHub repositories are supported. There is no user-scoped GitHub auth; private-repo support would require a deliberately designed OAuth flow.
- Without valid Upstash REST configuration, the analysis rate limiter remains per-instance and best-effort. Durable cross-instance quotas require the configured shared-store path or an external WAF/API-gateway rule.
- Blob-store retention relies on the cron sweep (
GET /api/cron/cleanupfor Vercel Cron, or authenticatedPOSTfor an operator scheduler); setCRON_SECRETin production.
Acceptance criteria: Zip with ../../etc/passwd does not write outside extract dir; analyzer never executes repo code; a private repo cannot be read via server credentials; JSON zipRef cannot analyze an arbitrary server path; corrupt stored report JSON is not served (see src/lib/reportSchema.test.ts, src/app/api/reports/reports-api.integration.test.ts).
| Strategy | Description |
|---|---|
| Staged analysis | Inventory, language packs, scoring, and report assembly run in sequence inside one request; the UI shows one honest loading state rather than fabricated stage progress |
| Same-commit cache | Complete public-GitHub reports may be reused by owner/repo@commit, analysis intent, and report version within the configured report TTL; invalid or partial records are not reused |
| Timeouts | GitHub API request: 15s; archive download: 60s; full analysis: 120s |
| Graceful degradation | After indexing, an expired analysis deadline returns a validated partial: true report with a Candidate Brief and warnings (see src/analyzer/index.ts) |
- Parsers: TypeScript Compiler API extraction, Python import-statement scanning, and Java import and same-package reference extraction.
- Language detection: extension → language.
- Scoring:
StartHereScoreandRiskScorewith mock inputs. - Path traversal: zip extraction rejects malicious paths.
- Full analyze flow: POST
/api/analyzewith a multipart fixture or callanalyzeRepository({ zipRef })internally, then assert either{ reportId, persisted: true }plus a saved artifact or{ reportId, report, persisted: false }plus a renderable inline report. - Public API rejects JSON
zipRefwith400 INVALID_INPUT. - Error mapping: invalid payloads/content types → documented
INVALID_INPUT,ZIP_INVALID,REPO_TOO_LARGE,TIMEOUT. - Stored report validation: corrupt JSON →
404on GET.
| Fixture | Description | Path |
|---|---|---|
fixtures/repo-ts |
Small TS repo (5 files): index, utils, 1 test | fixtures/repo-ts/ |
fixtures/repo-python |
Small Python repo: main, utils, test | fixtures/repo-python/ |
fixtures/repo-java |
Small Java repo: Main, Util, Test | fixtures/repo-java/ |
fixtures/repo-java-maven |
Maven Java app with tests | fixtures/repo-java-maven/ |
fixtures/repo-fastapi |
FastAPI service with tests | fixtures/repo-fastapi/ |
fixtures/repo-node-api |
Express API with package scripts | fixtures/repo-node-api/ |
fixtures/repo-monorepo |
npm workspaces monorepo | fixtures/repo-monorepo/ |
fixtures/repo-docs-only |
README-only repo (no source) | fixtures/repo-docs-only/ |
fixtures/repo-no-readme |
Minimal JS without README | fixtures/repo-no-readme/ |
fixtures/repo-docs-dedup |
Root + nested READMEs, byte-identical duplicate, distinct package README, duplicate run commands | fixtures/repo-docs-dedup/ |
Documentation-discovery edge cases (whitespace-only duplicates, similar-but-different docs, no-README repos, monorepo package docs) are additionally covered by programmatic temp-dir fixtures in src/analyzer/docs.test.ts.
- Folder Map tab: Renders non-empty tree for fixture.
- Architecture tab: Renders at least one node and edge for TS fixture.
- Start Here tab: Root README and entrypoint in list.
- Danger Zones tab: At least one file with score and breakdown.
- Run & Contribute tab: At least run commands or contribute signals.
- Export: PDF and PNG produce valid files for saved and inline reports; Markdown succeeds for a saved report and is unavailable with accurate guidance for an inline report.
For a minimal TS repo:
src/index.ts -> src/utils.ts
src/index.ts -> src/api/client.ts
src/api/client.ts -> src/utils.ts
Generated Mermaid (for markdown artifact rendering, not runtime UI):
flowchart LR
A["src/index.ts"]
B["src/utils.ts"]
C["src/api/client.ts"]
A --> B
A --> C
C --> B
{
"repo_metadata": {
"name": "tiny-app",
"url": "https://github.com/example/tiny-app",
"branch": "main",
"clone_hash": "abc123",
"analyzed_at": "2025-02-14T12:00:00.000Z"
},
"folder_map": {
"path": ".",
"type": "dir",
"children": [
{
"path": "README.md",
"type": "file"
},
{
"path": "src",
"type": "dir",
"children": [
{ "path": "src/index.ts", "type": "file" },
{ "path": "src/utils.ts", "type": "file" }
]
}
]
},
"architecture": {
"nodes": [
{ "id": "src/index.ts", "label": "index.ts" },
{ "id": "src/utils.ts", "label": "utils.ts" }
],
"edges": [
{ "from": "src/index.ts", "to": "src/utils.ts", "type": "import" }
]
},
"start_here": [
{ "path": "README.md", "score": 100, "explanation": "Root README" },
{ "path": "src/index.ts", "score": 85, "explanation": "Main entrypoint (package.json main)" }
],
"danger_zones": [
{
"path": "src/utils.ts",
"score": 72,
"breakdown": "High fan-in (1), no nearby tests",
"metrics": { "fan_in": 1, "fan_out": 0, "complexity": 5, "test_proximity": 0 }
}
],
"run_commands": [
{ "source": "package.json", "command": "npm run dev", "description": "Start dev server" }
],
"contribute_signals": {
"key_docs": ["README.md"],
"ci_configs": [".github/workflows/ci.yml"]
},
"warnings": []
}MVP items are shipped. For current status see CHANGELOG.md; for planned work see roadmap.md. When changing enforced behavior, update this spec and add an ADR under docs/adr/ when the decision is security- or policy-sensitive.
- Account-based or database-backed report ownership
- Persistent report storage beyond local JSON files or optional private Vercel Blob JSON
- Security vulnerability scanning
- Executing or profiling repository code
- Full AST parsing for all languages (use heuristics where AST is costly)
- Real-time collaboration
- Private GitHub repo support (MVP: public only)