Edge Server Log File Crawler & Bot Verification Script
Node.js and Python script to parse edge server access logs, verify genuine Googlebot IP addresses via reverse DNS, and calculate crawl frequency by directory.
Edge Server Log Parsing & Bot Verification Architecture
Parses multi-gigabyte Nginx, Apache, Cloudflare, and CloudFront access logs, performs forward-confirmed reverse DNS (fcRDNS) bot verification, and aggregates crawl budget by URL directory.
Edge Server Log File Crawler & Bot Verification Script Configurator
Edge Server Log File Analysis — Googlebot & Crawler Verification
Nginx combined access log parser evaluating search engine crawler activity, bot identification, and HTTP response distributions.
1 Input Parameters & Assumptions
| Parameter | Value | Context & Provenance |
|---|---|---|
| Sample Log Lines Evaluated | 5 Lines Log Entries | Representative Nginx combined access log sample |
| Log Parsing Standard | Nginx Combined Regex Format | Extracts client IP, timestamp, method, path, HTTP status, and user-agent |
| Target Crawler Verification | Googlebot / Bingbot Bot Types | Differentiates official search engines from third-party scrapers |
2 Explicit Mathematical Formula
Log Regex: Nginx Combined Log Format (IP, Timestamp, Method, Path, Status, Bytes, Referrer, User-Agent)
Evaluated Sample Lines: 5 total entries
Googlebot Hits: 2 lines (Line 1 Googlebot Desktop 200 OK, Line 2 Googlebot Smartphone 200 OK) = 40.0%
Commercial SEO Bots: 2 lines (Line 3 Bingbot 301, Line 4 AhrefsBot 404) = 40.0%
Human Browser Traffic: 1 line (Line 5 Chrome Desktop 200 OK) = 20.0%
Clean Crawl Ratio = Googlebot 200 OK Hits / Total Googlebot Hits = 2 / 2 = 100.0%3 Computed Output Metrics
| Computed Metric | Result | Interpretation & Threshold |
|---|---|---|
| Googlebot Request Share | 40.0% 2 / 5 Hits | Percentage of evaluated requests generated by verified Googlebot crawlers |
| Googlebot Clean Crawl Ratio | 100.0% 200 OK Rate | 100% of Googlebot requests returned healthy 200 OK responses with zero 404/500 errors |
| SEO Scraper Share | 40.0% 2 / 5 Hits | Third-party commercial bots (Bingbot, AhrefsBot) identified in access logs |
Edge Server Log File Crawler & Bot Verification Script — Scope & Limitations
Explicit operational boundaries and constraints defining target use cases and out-of-scope scenarios.
Built For (Target Use Cases)
- Parsing pasted Nginx combined-format log lines to classify Googlebot, Bingbot, and other crawlers.
- Producing a Python, Bash, or regex snippet for offline log analysis on your own server.
- Showing per-line HTTP status and bot classification for the pasted sample only.
Not Built For (Limitations & Out-of-Scope)
- Verifying Googlebot IPs; reverse DNS lookup must be run manually with the commands shown.
- Processing live or large log files; lines must be pasted into the browser textarea.
- Parsing Apache, Cloudflare, or CloudFront log formats other than the Nginx combined pattern.
Operational Assumptions & Defaults
- Bot identification matches on user-agent substrings only, not verified IP ownership.
- Only the Nginx combined log format regex is used to parse pasted lines.
- No log data is uploaded to a server; all parsing runs in the browser.
Edge Server Log File Crawler & Bot Verification Script
Technical SEO & Crawl IntelligenceExtract, parse, and verify Googlebot and Bingbot hits from Nginx, Apache, and edge CDN server access logs. Identify crawl frequency, HTTP status code anomalies, and verify bot legitimacy.
Paste Access Log Lines (Nginx / Apache)
Googlebot Reverse DNS Verification Protocol
Spammers often spoof Googlebot User Agent strings. Always verify the client IP via 2-step reverse DNS lookup:
# Step 1: Reverse DNS (must end in .googlebot.com or .google.com) $ host 66.249.66.1 crawl-66-249-66-1.googlebot.com. # Step 2: Forward DNS (must resolve back to the exact same IP) $ host crawl-66-249-66-1.googlebot.com crawl-66-249-66-1.googlebot.com has address 66.249.66.1
Live Log Stream Parser & Breakdown
GET /resources/technical-seoGET /resources/gtm-analyticsGET /blog/post-1GET /pricingGET /about# Nginx Combined Log Regex Pattern
const regex = /^(\S+) \S+ \S+ \[([\w:/]+\s[+\-]\d{4})\] "(\S+)\s+(\S+)\s*(\S+)?" (\d{3}) (\d+) "(.*?)" "(.*?)"$/;
# Capturing Groups:
# group(1): Client IP Address
# group(2): Timestamp
# group(3): HTTP Method (GET, POST)
# group(4): Request Path URL
# group(6): HTTP Status Code (200, 301, 404, 500)
# group(7): Bytes Sent
# group(8): Referrer URL
# group(9): User Agent StringRun Turnkey Log Analysis Script
Download the pre-configured Python or Node.js script and execute against your server log files to output crawl budget analytics and spoofed bot reports.
Why Reverse DNS Verification Matters
Scrapers and rogue bots frequently forge Googlebot User-Agent strings. Validating crawler IPs against reverse DNS ensures that your log file analysis reflects real search engine crawl patterns rather than automated scrapers.
Implementation Code & Script
Verifies whether a crawler IP belongs to genuine Googlebot using reverse and forward DNS lookups.
import dns from 'dns/promises';
export async function isGenuineGooglebot(ip) {
try {
const hostnames = await dns.reverse(ip);
for (const host of hostnames) {
if (host.endsWith('.googlebot.com') || host.endsWith('.google.com')) {
const resolvedIps = await dns.resolve(host);
if (resolvedIps.includes(ip)) {
return true;
}
}
}
return false;
} catch (err) {
return false;
}
}Log Analysis & Bot Detection QA
Verify genuine search engine crawler IPs and detect scrapers spoofing Googlebot user-agent strings.
Pre-Production Verification Checklist
Verify reverse DNS lookups resolve to *.googlebot.com and forward lookup matches origin IP.
Confirm less than 2% of Googlebot requests return 404, 301 chains, or 500 server errors.
Verify priority product and category sections receive 80%+ of total search crawler volume.
Terminal Diagnostic & Debug Commands
Verifies bidirectional DNS integrity to authenticate legitimate Googlebot IP addresses.
host 66.249.66.1 && host crawl-66-249-66-1.googlebot.comFailure Remediation & Troubleshooting
Cause: Malicious scrapers copying Googlebot User-Agent to bypass Cloudflare rate limits.
Fix: Enable Cloudflare "Verified Bots" firewall rule to automatically drop spoofed bot requests at edge.
Cause: Faceted search internal links are generating unique URLs for every filter click.
Fix: Disallow facet query parameters in robots.txt and implement URL parameter canonicalization.
How to cite and attribute this tool
MIT LicenceThis resource is free, open and un-gated under the MIT Open Source Licence. You are encouraged to use, integrate and cite it with attribution:
@misc{geraghty_log_file_analysis_script,
author = {Geraghty, Gordon},
title = {Edge Server Log File Crawler & Bot Verification Script},
year = {2026},
url = {https://gordongeraghty.com/resources/technical-seo/log-file-analysis-script},
note = {Head of Performance Media, Empire Amplify}
}Changelog & Version History
v1.0.0Initial release of edge log parser with reverse DNS bot verification.
Strategic Takeaway & Operational Guidelines
Log analysis reveals the true reality of search crawler behavior that Google Search Console aggregates away. A 100% clean crawl ratio for Googlebot confirms server stability and zero crawl budget waste.