A Go application that extracts and lists all container images used in Dockerfiles across GitHub repositories.
This tool scans Dockerfiles from specified GitHub repositories and generates a JSON output that maps repositories to their container image dependencies.
- Authenticates with GitHub (optional for higher API limits)
- Scans Dockerfiles in the specified repositories
- Extracts all container images (FROM statements) from Dockerfiles
- Generates JSON output mapping repositories to their container images
{
"data": {
"owner/repo1": {
"Dockerfile": ["nginx:1.19", "node:14"],
"services/Dockerfile": ["python:3.9"]
},
"owner/repo2": {
"Dockerfile.prod": ["golang:1.16"]
}
}
}- GitHub access token with repo scope (optional, for private repository access and higher API rate limits)
The application requires the following environment variables:
REPOSITORY_LIST_URL: URL to the text file containing repository sourcesGITHUB_ACCESS_TOKEN: GitHub access token for API authentication (optional)
git clone https://github.com/umutondersu/dockerfile-sources
cd dockerfile-sources
go build ./cmd/dockerfile-sources
# dockerfile-sourcesThe application can be deployed as a Kubernetes job. The necessary configuration files and helper scripts are provided in the k8s directory.
- Minikube
- kubectl
- Docker
k8s/job.yaml: Kubernetes job configuration with resource limits and timeoutsk8s/secret.yaml: Template for usingGITHUB_ACCESS_TOKENk8s/scripts/: Helper scripts for managing the job
The following scripts are provided to simplify deployment and monitoring:
startjob.sh: Handles Minikube startup, builds Docker image, and deploys the jobgetjoblogs.sh: Monitors job logs by automatically finding the podendjob.sh: Cleans up job resources
- Make the scripts executable:
chmod +x k8s/scripts/*.sh- Start the job:
./k8s/scripts/startjob.sh- Monitor logs (in a separate terminal):
./k8s/scripts/getjoblogs.sh- Clean up when done:
./k8s/scripts/endjob.shThe application is organized into several packages, each with a specific responsibility:
- Handles parsing of repository source data
- Contains the
Sourcestruct that represents a GitHub repository with owner, repo name, and commit SHA - Implements HTTP response fetching for repository list URL
- Uses regex pattern matching to extract repository information
- Manages GitHub API interactions through a custom client
- Responsible for scanning repositories for Dockerfile presence
- Extracts container image information from Dockerfiles
- The
DockerFilestruct maintains the relationship between sources and their container images - Implements concurrent processing of multiple repositories
- Handles the conversion of Dockerfile data into the final JSON format
- Implements the
OutputDatastructure that maps repositories to their Dockerfile paths and images - Provides JSON generation functionality with proper error handling
- Ensures consistent output format as shown in the example above
- Contains the main application entry point
- Orchestrates the workflow:
- Reads environment variables
- Fetches repository list
- Initializes GitHub client
- Processes Dockerfiles
- Generates and outputs JSON
- Application reads repository list from provided URL
- For each repository:
- Scans for Dockerfile presence
- Extracts FROM statements to identify container images
- Uses goroutines and channels for efficient file content retrieval
- Implements wait groups to ensure all processing completes
- Aggregates results into a structured JSON output
- Prints final JSON to stdout
- Concurrent processing of Dockerfile content
- Channel-based communication for efficient data transfer
- Wait group synchronization for parallel operations
- Optimized memory usage through streaming processing
- Comprehensive error checking at each step
- Validation of input data format
- Operation timeout control (default: 5 minutes)
- Graceful failure handling for GitHub API rate limits and network errors
- Proper error propagation through the application
- Robust error handling system with:
- Centralized error processing through
handleGitHubResponseError - Smart retry mechanism with exponential backoff for transient failures
- Standardized error patterns across all GitHub API calls
- Detailed error types for better debugging:
- Maximum retry duration of 30 seconds with 100ms initial interval
- Automatic backoff for retryable errors (500+ status codes)
- Concurrent operation error capture and reporting
- Centralized error processing through
- Consistent exit codes for different error scenarios
- How ghdocker.getFileContent handles errors in concurrency: Indicating incomplete results, non-silent failure and Aggregating errors
-
Use Full Concurrency
-
Handle context cancellation of parent inside GetDockerFiles
-
For Concurrent process use a worker pool to prevent resource exhaustion
- getFileContent doesn't need to return the file content since its already sending it to the channel (It was like this for testing and pre-concurrency)
- Redundant use of methods while basic functions would suffice
- Sync backoff algorithm with Github API's rate limits
- Image Name variable shouldn't have a tag in k8s/startjob.sh
- Use WithAuthToken() for Initializing the github client instead of importing auth2 (optional)