Skip to content

Integrating ClickHouse with GitHub

The GitHub registry item installs one pipeline with independently selected resource streams and separate raw ClickHouse destinations.

Terminal window
bunx chkit add github --with-tests
bunx chkit check
bunx chkit generate --name add_github
bunx chkit migrate --apply
bunx chkit ingest run --tag provider:github

Set GITHUB_TOKEN before the ingestion run. For private repositories, the supported setup grants repository access with Issues: read, Pull requests: read, and Contents: read, including GraphQL commit reads. A field-level permission minimum has not been verified. Edit githubConfig.repositories, source identity, and reader settings in src/integrations/github/config.ts. Raw tables are declared in their sources/ modules before schema discovery and migration. Native JSON requires ClickHouse 25.3 or later; schema imports do not call GitHub.

ResourceDefault ClickHouse tableRecords syncedAPI reference
Issues (issues)github_issues_rawOnly ordinary issues, with complete provider issue fields and repository contextGET/repos/{owner}/{repo}/issues
Issue comments (issue_comments)github_issue_comments_rawOne raw comment per row for ordinary issues; independently discovers all issue parents and filters child updatesGET/repos/{owner}/{repo}/issues, GET/repos/{owner}/{repo}/issues/{issue_number}/comments
Pull requests (pull_requests)github_pull_requests_rawPull-request metadata from its own collection, with repository contextGET/repos/{owner}/{repo}/pulls
Pull request conversation comments (pull_request_comments)github_pull_request_comments_rawOne raw conversation comment per row with its pull number; independently discovers all pull-request parents and filters child updatesGET/repos/{owner}/{repo}/pulls, GET/repos/{owner}/{repo}/issues/{pull_number}/comments
Pull request review comments (pull_request_review_comments)github_pull_request_review_comments_rawRaw diff-review comments with repository and pull-number context, selected by their own update timesGET/repos/{owner}/{repo}/pulls/comments
Pull request reviews (pull_request_reviews)github_pull_request_reviews_rawOne raw review per row with repository and pull number; discovers its own pull-request parentsGET/repos/{owner}/{repo}/pulls, GET/repos/{owner}/{repo}/pulls/{pull_number}/reviews
Pull request commits (pull_request_commits)github_pull_request_commits_rawCommit SHA and message only, with repository and pull number; paginates GraphQL commits after independent pull-request discoveryGET/repos/{owner}/{repo}/pulls, POST/graphql
Stargazers (stargazers)github_stargazers_rawRepository stargazers with starred_at and repository contextGET/repos/{owner}/{repo}/stargazers

Each configured repository contributes independent streams for issues, issue comments, pull requests, PR conversation comments, PR review comments, PR reviews, PR commits, and stargazers. Issues exclude pull requests. The streams share one installation pipeline and retain their own journal progress; stream order creates no dependency.

Child-resource streams discover their own parents through the API. Selecting PR comments or commits works even when the pull-requests stream has never run. Individual PRs and issues remain records within those resource streams.

Terminal window
bunx chkit ingest run --tag provider:github --tag repository:obsessiondb/chkit --tag resource:pull_requests
bunx chkit ingest run --tag provider:github --tag repository:obsessiondb/chkit --tag resource:pull_request_commits

Repeated tags use AND matching. The default pipeline runs streams sequentially with maxStreams: 1; filtered runs give each resource its own schedule and execution budget. createGitHubPipeline(config, deps) snapshots reader settings and injectable HTTP dependencies. Changing sourceId creates new stream identities; runtime reader settings do not redefine the exported raw tables.

Raw payloads use { repository, data }. Issue comments add issue_number; PR child rows add pull_number. Provider response fields stay under data. PR commits retain only data.sha and data.message, with repository/PR join keys. Their reader paginates the GraphQL commit connection, rejects repeated SHAs, and checks stable totals and distinct terminal counts before completing the stream. Changed files and patches are not fetched.

Join resources in ClickHouse through repository and issue/PR number. PR metadata lives in its own destination, and comments, reviews, and commits are separate rows. Conversation comments and diff-review comments remain distinct collections. Existing issue/stargazer provider-ID formats are preserved.

Issues and comments use independent updated-time windows with five minutes of overlap and a fixed local upper cutoff. Child comments on old issues and PRs remain eligible because parent discovery does not filter by parent update times. Review comments use the repository-wide update feed and retain their PR number. The checkpoint advances only after the selected stream’s complete window and destination writes succeed; failed windows replay from their lower bound. See issue filtering, issue comments, and review comments.

Pull requests, reviews, commits, and stargazers use fullSync(). Each invocation reads its selected collection from the beginning, and the library journals successful completion, including empty results. REST Link continuations and GraphQL cursors use paginate() for request retries and cycle detection; no mutable page offset or parent scan ledger is saved.

Deleted records, removed stars, and commits removed from PRs remain stored. Live pagination does not guarantee a source snapshot. Historical replay refreshes current observations outside the timestamp overlap:

Terminal window
bunx chkit ingest run --tag provider:github --tag resource:issue_comments --backfill reconcile-2026-10-05 --from 2008-01-01
bunx chkit ingest status --tag provider:github --json

Use a new backfill ID for each reconciliation. Full-sync resources ignore date bounds. A failed stream does not advance another stream’s progress. Schedule finite ingestion runs externally, with one process per ClickHouse target at a time.

Version 0.2.0 adds six raw destinations and changes payloads to { repository, data }. Existing issue and stargazer row IDs stay unchanged. Retained flat payloads may coexist with new envelopes until they are reread or migrated.

Existing combined-reader datasets may include PR rows in github_issues_raw and child arrays inside parent payloads. Those rows remain after the new streams run. Archive the legacy dataset or deliberately migrate classifications and child records into the separate destinations. Normal ingestion does not remove legacy PR rows from the issues table; application queries must account for retained shapes during migration.

Terminal window
bun test src/integrations/github/tests/basic.test.ts

Fixtures cover isolated resource selection, join keys, narrow commit data, recent comments on old parents, pagination, failures, and journal completion without GitHub credentials.

Version 0.2.0

  • Sync issues and editable comments through overlapping updated-time windows; advance each stream watermark only after its destination acknowledgement.
  • Separate issues, issue comments, pull requests, PR conversation comments, review comments, reviews, commits, and stargazers into independently selectable raw streams for every configured repository.
  • Store repository and parent join keys beside provider data; retain only commit SHA and message through paginated GraphQL commit connections, and leave joins to ClickHouse.
  • Use one installation pipeline with configuration snapshots, injectable requests, shared pagination, and fullSync where no reliable change filter exists; existing mixed or combined observations require deliberate data migration.
  • Consume the ingestion plugin's full pagination pages, preserving continuations and empty terminal pages with independent resource checkpoints.

Version 0.1.0

  • Introduce raw repository issues and stargazers ingestion.