Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 15 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,20 @@ jobs:
- name: Install dependencies
run: npm ci

# Lint — Biome's correctness + suspicious rules only. Style rules are
# mostly disabled to avoid a big churn day-one; the goal here is to
# catch real bugs (unused imports, shadowed variables, assign-in-cond)
# rather than enforce a specific house style.
- name: Lint
run: npm run lint

# Unit tests — Vitest covers the pure-logic modules (risk classifier,
# EU AI Act evaluator + scoring, AI compliance risk generation) with
# focused per-function assertions. Adds meaningful coverage beyond
# "did not throw" while still running in milliseconds.
- name: Unit tests
run: npm test

# Run the scanner against this repo itself. Catches scanner-side
# regressions end-to-end: scan rules, policy rendering, framework
# evaluation, and report generation all exercised on a real tree.
Expand All @@ -40,4 +54,4 @@ jobs:
# outage where an older stored manifest took /api/repos and the
# homepage down with an uncaught TypeError.
- name: Dashboard render smoke
run: npx tsx scripts/smoke-dashboard.ts
run: npm run smoke:dashboard
14 changes: 12 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ On every push or PR, the scanner produces:
| `security.txt` | `.well-known/` | RFC 9116 security contact file |
| `risk-assessment.md` | `.grc/` | Likelihood x impact matrix with framework mappings |
| `nist-csf-report.md` | `.grc/` | 18 NIST CSF controls with SOC 2 + ISO 27001 cross-mapping |
| `security-headers-report.md` | `.grc/` | Header status + copy-paste fix |
| `security-headers-report.md` | `.grc/` | Header status + starter-snippet fixes (CSP typically needs manual review) |
| `access-controls-report.md` | `.grc/` | Branch protection and auth findings |

Reports (`.grc/`) are gitignored and regenerated each scan. Policies (`docs/policies/`, `.well-known/`) are committed to your PR branch so they ship with your code.
Expand Down Expand Up @@ -208,7 +208,7 @@ Badge states:
- `fail NN%` — critical vulnerabilities, detected secrets, or very low compliance
- `not scanned` — the dashboard has no manifest for that repo/branch yet

GitHub's GitHub App badge UI is a separate static logo upload. See [docs/github-app-badge.md](docs/github-app-badge.md) for the distinction and setup steps.
GitHub's GitHub App badge UI is a separate static logo upload. See [docs/badges.md](docs/badges.md) for the distinction and setup steps.

## Scanner

Expand All @@ -234,6 +234,16 @@ Reports are written to `/path/to/repo/.grc/`. Policies are written to `/path/to/
- **TLS** - HTTPS enforcement, certificate expiry (live URL check)
- **AI Systems** - detects AI SDKs (OpenAI, Anthropic, Cohere, Gemini, HuggingFace, Mistral, Groq, LangChain, LlamaIndex, Vercel AI SDK), training libs (TensorFlow, PyTorch), vector DBs (Pinecone, Weaviate, ChromaDB, Qdrant), and outbound API calls. Supports Node (`package.json`), Python (`requirements.txt`, `pyproject.toml`), and monorepos.

### Language coverage

The scanner's Node/JavaScript path is the most mature — forms, endpoints, dependencies, secrets, tracking, and AI SDK detection all fully work on `.ts` / `.tsx` / `.js` / `.jsx` / `.mjs` / `.cjs` trees.

**Python** support is partial: `requirements.txt` and `pyproject.toml` are scanned for AI packages and third-party services, but form/endpoint/secret detection only has basic regex coverage. Flask, Django, and FastAPI idioms aren't specifically recognised yet.

**Go, Ruby, Java, Rust, PHP** — not meaningfully supported. Files are walked for secret regexes and outbound AI API URL patterns; nothing else. A repo in any of these languages will scan without erroring but the findings list will be sparse compared to a Node repo.

If you're running the scanner against a non-Node repo, expect partial signal and treat missing findings as absence of evidence, not evidence of absence.

## AI Enhancements (Optional)

Add to `.grc/config.yml`:
Expand Down
57 changes: 57 additions & 0 deletions biome.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
{
"$schema": "https://biomejs.dev/schemas/2.4.12/schema.json",
"files": {
"includes": [
"scanner/**/*.ts",
"dashboard/**/*.ts",
"scripts/**/*.ts",
"!node_modules/**",
"!dist/**",
"!.grc/**",
"!.well-known/**",
"!.wrangler/**"
]
},
"linter": {
"enabled": true,
"rules": {
"recommended": false,
"correctness": {
"noUnusedVariables": "error",
"noUnusedImports": "error",
"noUnusedPrivateClassMembers": "error",
"useExhaustiveDependencies": "off",
"noInvalidUseBeforeDeclaration": "error"
},
"suspicious": {
"noDoubleEquals": "error",
"noDuplicateCase": "error",
"noDuplicateObjectKeys": "error",
"noDuplicateParameters": "error",
"noConstEnum": "error",
"noAssignInExpressions": "error",
"noExplicitAny": "off",
"noConsole": "off",
"useAwait": "off"
},
"style": {
"noNonNullAssertion": "off",
"useTemplate": "off",
"useConst": "error",
"useSingleVarDeclarator": "off",
"noParameterAssign": "off"
},
"complexity": {
"noForEach": "off",
"noUselessConstructor": "off",
"useOptionalChain": "off"
},
"performance": {
"noDelete": "off"
}
}
},
"formatter": {
"enabled": false
}
}
2 changes: 1 addition & 1 deletion dashboard/worker.ts
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
import { Hono } from "hono";
import { parse } from "yaml";
import type { Manifest, AISystem } from "../scanner/types.js";
import type { Manifest, } from "../scanner/types.js";
import { evaluateFramework } from "../scanner/generators/framework-report.js";
import { evaluateEUAIAct, calcAIComplianceScore } from "../scanner/frameworks/eu-ai-act.js";
import { renderDashboard, renderRepoDetail, renderNistView, renderBranchComparison, renderTrendChart, renderAIComplianceView, renderInventoryView } from "./views/render.js";
Expand Down
145 changes: 67 additions & 78 deletions docs/architecture.md
Original file line number Diff line number Diff line change
@@ -1,92 +1,81 @@
# Architecture

## System Design: Hub and Spoke
## Hub and spoke

```
┌─────────────────────────────────────────┐
│ Central GRC Dashboard │
│ (aggregates everything) │
└──────────┬──────────┬──────────┬────────┘
│ │ │
Repo A Repo B Repo C
(GH Action) (GH Action) (GH Action)
scans on scans on scans on
PR/deploy PR/deploy PR/deploy
┌──────────────────────────┐
│ Cloudflare Worker │
│ (dashboard + KV) │
└──────────┬───────────────┘
│ POST /api/report (OIDC-authed)
│
┌─────────────┼─────────────┐
│ │ │
Repo A Repo B Repo C
GitHub Action GitHub Action GitHub Action
writes .grc/ writes .grc/ writes .grc/
commits policy commits policy commits policy
```

Each repo:
- Runs the same reusable GitHub Action
- Generates a local compliance badge and manifest file committed to the repo
- POSTs the manifest to the central dashboard API
Every consuming repo runs the same composite action (`action.yml`). Each run:

The central dashboard:
- Receives and stores manifests from all repos
- Provides the org-wide view, trends, and reporting
1. Scans the repo tree (Node source + `package.json` + `requirements.txt` + `pyproject.toml`) and the live URL if configured.
2. Writes a YAML manifest to `.grc/manifest.yml` and generated policy markdown to `docs/policies/` + `/.well-known/security.txt`.
3. On PRs, commits generated policies back to the PR branch (attributed to `grc-bot`).
4. Mints a short-lived GitHub OIDC JWT and POSTs the manifest to the dashboard's `/api/report`.

## Single Source of Truth Principle
The dashboard (Hono on Cloudflare Workers) verifies the JWT against GitHub's JWKS, stores the manifest in KV keyed by `manifest:<repo>:<branch>`, and renders HTML views with HTMX.

## Single source of truth

The scan is the authority.

```
❌ Traditional (drift-prone):
❌ Drift-prone:
Lawyer writes policy → hope code matches → audit finds gaps

✅ Our approach (compliance-as-code):
Code is scanned → scan produces facts → facts generate policy
→ facts feed dashboard
✅ Compliance-as-code:
Code scanned → scan produces facts → facts generate policy
→ facts feed dashboard
→ facts feed framework scoring
```

The scan is the authority. The privacy policy, ToS, and dashboard status are all derivatives of what the scan found. If the scan says "this repo collects email via a form and sends it through Resend," then:
- The privacy policy gets a section about email collection and Resend as a processor
- The dashboard shows "data collection: email (processor: Resend)" as a tracked item

## GitHub Action Flow

The reusable GitHub Action runs in its own container and can scan ANY repo regardless of language:

**Outputs:**
1. `manifest.yml` — structured compliance data (committed to repo)
2. POST to dashboard API — feeds the central dashboard
3. Compliance badge SVG — visual status indicator
4. Generated policies — only for repos serving public-facing sites

## Tech Stack

- **GitHub Action**: Reusable workflow (`.github/workflows/grc-scan.yml`)
- **Scanner**: Node.js script using AST parsing + regex
- **Dashboard API**: Express endpoint (or separate service)
- **Dashboard UI**: HTMX (matches joeeftekhari.com stack)
- **Storage**: Postgres on Digital Ocean droplet (or JSON files to start)

## Build Tiers

### Tier 1 (Start Here)
- Scanner detects: data collection, security headers, dependencies, secrets, TLS, security.txt
- Outputs: manifest.yml, generated policies (privacy policy, ToS, security.txt, vulnerability disclosure)
- Dashboard: checklist view per repo

### Tier 2 (After Tier 1 Works)
- Add: framework mapping (NIST CSF controls → scan results)
- Add: branch comparison (main vs feature branches)
- Add: trend tracking over time
- Dashboard: framework compliance percentages

### Tier 3 (Portfolio Showstopper)
- Add: audit evidence export (PDF/ZIP per framework)
- Add: AI-powered gap analysis ("you're missing X for SOC 2")
- Add: remediation suggestions with auto-fix PRs
- Dashboard: auditor-ready report generation

## Where AI Fits In

| Task | Deterministic Scan | AI Layer |
|---|---|---|
| "Is there a form?" | Regex for `<form`, POST routes | — |
| "What data does it collect?" | Parse input names | Classify as PII vs non-PII |
| "Is this a new third-party service?" | Check imports/API calls | Determine if it's a data processor |
| "Is the policy still accurate?" | Diff manifest vs last policy | Explain what changed in plain English |
| "What should we do about this?" | — | Generate remediation steps |

Additional AI opportunities:
- LLM reviews code diffs for new data collection patterns humans might miss
- Auto-classifies data types (PII vs non-PII)
- Generates remediation suggestions when a check fails
- Summarizes compliance posture changes in plain English for non-technical stakeholders
If the scan finds that a repo collects email via a form and sends it through Resend, downstream outputs follow automatically: the privacy policy names Resend as a processor, the dashboard's data-collection row lists "email (processor: Resend)", and the NIST CSF check for data inventory flips to `pass` based on that evidence.

## Tech stack

- **Scanner:** Node 20, TypeScript via `tsx` at runtime. No build step.
- **Scan rules:** `scanner/rules/*.ts` — one file per concept (forms, endpoints, secrets, dependencies, access controls, AI systems, …). Each returns structured findings.
- **Policy templates:** Handlebars (`.hbs`) in `scanner/templates/`. Rendered to markdown by `scanner/render.ts`.
- **Reports:** `scanner/generators/*.ts` — markdown output per framework / concern (NIST CSF, EU AI Act, risk assessment, security headers, access controls).
- **Dashboard:** Hono on Cloudflare Workers. Inlined HTML + Press Start 2P / JetBrains Mono + HTMX for tab navigation.
- **Storage:** Cloudflare KV, two key shapes: `manifest:<repo>:<branch>` for current state, `history:<repo>` for trend data.
- **Auth:** GitHub OIDC on `POST /api/report`. No shared secrets; see `dashboard/auth.ts`.

## Where AI fits in

The AI enhancement layer (Phase 4) is optional — the scanner works fully without it. When enabled and given an API key, an LLM refines PII classification, rewrites risk narratives in plain English, and generates gap analyses with prioritized recommendations.

The EU AI Act detection + risk classification (Phase 8) is a separate thing: purely deterministic scanning of consuming repos for AI SDK imports, framework imports, and outbound AI API URLs. No LLM involvement in that path.

## What each folder is for

| Folder | Purpose |
|---|---|
| `scanner/rules/` | Per-concept scan rules producing structured findings |
| `scanner/templates/` | Handlebars policy templates |
| `scanner/generators/` | Markdown report generators |
| `scanner/frameworks/` | Framework definitions (NIST CSF, EU AI Act) and cross-maps |
| `scanner/ai/` | Optional LLM enhancement layer |
| `dashboard/` | Cloudflare Worker + render functions |
| `dashboard/views/render.ts` | All HTML rendering — server-rendered, HTMX for tab swaps |
| `scripts/` | Standalone tsx utilities (smoke tests, one-off maintenance) |
| `docs/` | Reference documentation (this file, checklist, GRC fundamentals, badges) |

## Not covered here

- **What the scanner detects** — see the "What It Scans" section in the [README](../README.md).
- **How to set up a fork** — see [README § Setup](../README.md#setup).
- **The manifest schema** — `scanner/types.ts` is the authoritative source. The TypeScript types are the schema.
- **Roadmap** — [implementation-checklist.md](implementation-checklist.md).
- **How to extend the scanner** — [CONTRIBUTING.md](../CONTRIBUTING.md).
File renamed without changes.
100 changes: 0 additions & 100 deletions docs/compliance-scope.md

This file was deleted.

Loading
Loading