About this workflow
Scrape Today's GitHub Trending Repositories Workflow Analysis
This n8n workflow template is designed to automatically scrape and extract structured data from the GitHub Trending page (https://github.com/trending), specifically targeting the top repositories featured for the current day. It leverages HTML parsing capabilities to navigate GitHub’s public HTML structure, isolate relevant repository elements, and transform raw scraped content into a clean, usable list of repository details including author, name, description, and direct URL.
What the Workflow Does
The workflow initiates manually, fetches the HTML content of GitHub’s trending page, isolates the main container holding repository listings, extracts individual repository cards, parses key metadata (such as repository path, programming language, and description), and finally formats each entry into a standardized JSON object with enriched fields like author, title, url, and timestamped created_at. The output is a list of up to 25 trending repositories (as displayed by GitHub), ready for downstream use such as database insertion, notification alerts, or integration with other tools.
Key Features and Capabilities
- Web Scraping Without APIs: Bypasses the need for GitHub API tokens by directly scraping public HTML.
- Dynamic Data Extraction: Uses CSS selectors to reliably target repository elements even as GitHub’s layout evolves moderately.
- Structured Output: Converts unstructured HTML into consistent JSON records with normalized fields.
- Manual Trigger: Designed for on-demand execution via the “Test workflow” button, ideal for scheduled or ad-hoc runs.
- Text Cleaning: Includes options to trim whitespace and clean up extracted text for better readability and consistency.
- URL Construction: Dynamically builds valid GitHub repository URLs from scraped paths.
Main Nodes and Their Purposes
-
When clicking ‘Test workflow’ (
manualTrigger)- Serves as the entry point; triggers the workflow manually during testing or execution.
-
Request to Github Trend (
httpRequest)- Sends an HTTP GET request to
https://github.com/trendingto retrieve the raw HTML of the trending page.
- Sends an HTTP GET request to
-
Extract Box (
htmlnode)- Parses the full HTML response and extracts the content of the first
<div class="Box">, which contains the main list of trending repositories.
- Parses the full HTML response and extracts the content of the first
-
Extract all repositories (
htmlnode)- Processes the extracted "Box" HTML and isolates each
<article class="Box-row">element (representing individual repositories) into an array of HTML snippets usingreturnArray: true.
- Processes the extracted "Box" HTML and isolates each
-
Turn to a list (
splitOutnode)- Splits the array of repository HTML snippets into individual items, enabling item-by-item processing in subsequent nodes.
-
Extract repository data (
htmlnode)- For each repository HTML snippet, extracts three key pieces of information:
repository: The author/repo path from the<a class="Link">element.language: Programming language from<span class="d-inline-block">.description: Brief project description from the<p>tag.
- For each repository HTML snippet, extracts three key pieces of information:
-
Set Result Variables (
setnode)- Transforms and enriches the extracted data:
- Splits
repositoryintoauthorandtitle. - Constructs a valid
urlusing the repository path. - Adds a
created_attimestamp using{{$now}}. - Preserves the original
description. - Retains all other fields via
includeOtherFields: "=".
- Splits
- Transforms and enriches the extracted data:
Use Cases and Benefits
- Developer Research: Quickly discover new or rising open-source projects.
- Automated Monitoring: Track daily trends for competitive analysis or inspiration.
- Content Aggregation: Feed trending repos into dashboards, newsletters, or internal tools.
- Educational Purposes: Learn web scraping techniques using n8n’s visual automation.
- Low-Cost Alternative: Avoid rate limits or authentication requirements of the official GitHub API for public data.
Step-by-Step Workflow Logic
- Trigger: User clicks “Test workflow,” activating the manual trigger.
- Fetch Page: The
httpRequestnode retrieves the full HTML ofhttps://github.com/trending. - Isolate Container: The
Extract Boxnode usesdiv.Boxto capture the primary trending section. - Extract Repository Cards: The
Extract all repositoriesnode selects allarticle.Box-rowelements and returns them as an array of HTML strings under therepositoriesfield. - Split into Items: The
Turn to a listnode converts therepositoriesarray into individual workflow items for parallel or sequential processing. - Parse Metadata: For each repository item, the
Extract repository datanode applies CSS selectors to pull out:- Repository path (e.g.,
"n8n-io/n8n") - Language badge text
- Description paragraph
- Repository path (e.g.,
- Enrich & Standardize: The
Set Result Variablesnode:- Splits the repository path into
authorandtitle - Builds a clickable GitHub URL
- Adds current timestamp
- Cleans and assigns all fields into a uniform structure
- Splits the repository path into
- Output: The final output is a list of structured repository objects, ready for export, storage, or further automation.
This workflow exemplifies efficient, no-code web scraping using n8n’s native HTML parsing and data manipulation nodes, providing actionable insights from publicly available GitHub data.
How to use: download the JSON, then in n8n choose “Import from File”.