Skip to content

Edge Server Log File Crawler & Bot Verification Script

Node.js and Python script to parse edge server access logs, verify genuine Googlebot IP addresses via reverse DNS, and calculate crawl frequency by directory.

By Gordon Geraghty·MIT Licence·Updated: 24 September 2026·ADVANCED
01 Prerequisites & Architecture
Stage 01 Architecture

Edge Server Log Parsing & Bot Verification Architecture

Parses multi-gigabyte Nginx, Apache, Cloudflare, and CloudFront access logs, performs forward-confirmed reverse DNS (fcRDNS) bot verification, and aggregates crawl budget by URL directory.

Difficulty:Advanced Architecture
Time:25–35 mins
Required Access & Permissions:
Server Access Log Read PermissionsPython 3.9+ or Node.js 18+ Runtime
STEP 01Nginx / Cloudflare / CloudFront
Access Log StreamRaw server hits (.log / .gz / JSON)
STEP 02Python / Node.js Streaming Engine
Regex Format ParserExtracts IP, timestamp, method, URI, status, UA
STEP 03DNS Reverse Lookup (fcRDNS)
fcRDNS Bot VerificationValidates genuine Googlebot / Bingbot / AI scrapers
STEP 04CSV / JSON Report Export
Crawl Budget MatrixStatus code distribution & crawl frequency charts
02 Interactive Configurator

Edge Server Log File Crawler & Bot Verification Script Configurator

Worked Example · Deterministic Calculation

Edge Server Log File Analysis — Googlebot & Crawler Verification

Nginx combined access log parser evaluating search engine crawler activity, bot identification, and HTTP response distributions.

1 Input Parameters & Assumptions

ParameterValueContext & Provenance
Sample Log Lines Evaluated5 Lines Log EntriesRepresentative Nginx combined access log sample
Log Parsing StandardNginx Combined Regex FormatExtracts client IP, timestamp, method, path, HTTP status, and user-agent
Target Crawler VerificationGooglebot / Bingbot Bot TypesDifferentiates official search engines from third-party scrapers

2 Explicit Mathematical Formula

Log Regex: Nginx Combined Log Format (IP, Timestamp, Method, Path, Status, Bytes, Referrer, User-Agent)
Evaluated Sample Lines: 5 total entries
Googlebot Hits: 2 lines (Line 1 Googlebot Desktop 200 OK, Line 2 Googlebot Smartphone 200 OK) = 40.0%
Commercial SEO Bots: 2 lines (Line 3 Bingbot 301, Line 4 AhrefsBot 404) = 40.0%
Human Browser Traffic: 1 line (Line 5 Chrome Desktop 200 OK) = 20.0%
Clean Crawl Ratio = Googlebot 200 OK Hits / Total Googlebot Hits = 2 / 2 = 100.0%

3 Computed Output Metrics

Googlebot Request Share40.0%2 / 5 HitsPercentage of evaluated requests generated by verified Googlebot crawlers
Googlebot Clean Crawl Ratio100.0%200 OK Rate100% of Googlebot requests returned healthy 200 OK responses with zero 404/500 errors
Computed MetricResultInterpretation & Threshold
Googlebot Request Share40.0% 2 / 5 HitsPercentage of evaluated requests generated by verified Googlebot crawlers
Googlebot Clean Crawl Ratio100.0% 200 OK Rate100% of Googlebot requests returned healthy 200 OK responses with zero 404/500 errors
SEO Scraper Share40.0% 2 / 5 HitsThird-party commercial bots (Bingbot, AhrefsBot) identified in access logs

Strategic Takeaway & Operational Guidelines

Log analysis reveals the true reality of search crawler behavior that Google Search Console aggregates away. A 100% clean crawl ratio for Googlebot confirms server stability and zero crawl budget waste.

INSTRUMENT BOUNDARIES

Edge Server Log File Crawler & Bot Verification Script — Scope & Limitations

Explicit operational boundaries and constraints defining target use cases and out-of-scope scenarios.

Built For (Target Use Cases)

  • Parsing pasted Nginx combined-format log lines to classify Googlebot, Bingbot, and other crawlers.
  • Producing a Python, Bash, or regex snippet for offline log analysis on your own server.
  • Showing per-line HTTP status and bot classification for the pasted sample only.

Not Built For (Limitations & Out-of-Scope)

  • Verifying Googlebot IPs; reverse DNS lookup must be run manually with the commands shown.
  • Processing live or large log files; lines must be pasted into the browser textarea.
  • Parsing Apache, Cloudflare, or CloudFront log formats other than the Nginx combined pattern.

Operational Assumptions & Defaults

  • Bot identification matches on user-agent substrings only, not verified IP ownership.
  • Only the Nginx combined log format regex is used to parse pasted lines.
  • No log data is uploaded to a server; all parsing runs in the browser.

Edge Server Log File Crawler & Bot Verification Script

Technical SEO & Crawl Intelligence

Extract, parse, and verify Googlebot and Bingbot hits from Nginx, Apache, and edge CDN server access logs. Identify crawl frequency, HTTP status code anomalies, and verify bot legitimacy.

Paste Access Log Lines (Nginx / Apache)

Googlebot Reverse DNS Verification Protocol

Spammers often spoof Googlebot User Agent strings. Always verify the client IP via 2-step reverse DNS lookup:

# Step 1: Reverse DNS (must end in .googlebot.com or .google.com)
$ host 66.249.66.1
crawl-66-249-66-1.googlebot.com.

# Step 2: Forward DNS (must resolve back to the exact same IP)
$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1

Live Log Stream Parser & Breakdown

Googlebot Hits
2
of 5 total lines
200 OK Hits
3
1 4xx errors / 1 3xx
Googlebot DesktopGET /resources/technical-seo
200
IP: 66.249.66.1 • Time: 23/Aug/2026:14:32:10 +0000
Googlebot SmartphoneGET /resources/gtm-analytics
200
IP: 66.249.66.2 • Time: 23/Aug/2026:14:33:15 +0000
BingbotGET /blog/post-1
301
IP: 40.77.167.5 • Time: 23/Aug/2026:14:34:02 +0000
SEO Crawler (Ahrefs/Semrush)GET /pricing
404
IP: 54.36.148.1 • Time: 23/Aug/2026:14:35:40 +0000
Human / BrowserGET /about
200
IP: 192.0.2.100 • Time: 23/Aug/2026:14:36:12 +0000
# Nginx Combined Log Regex Pattern
const regex = /^(\S+) \S+ \S+ \[([\w:/]+\s[+\-]\d{4})\] "(\S+)\s+(\S+)\s*(\S+)?" (\d{3}) (\d+) "(.*?)" "(.*?)"$/;

# Capturing Groups:
# group(1): Client IP Address
# group(2): Timestamp
# group(3): HTTP Method (GET, POST)
# group(4): Request Path URL
# group(6): HTTP Status Code (200, 301, 404, 500)
# group(7): Bytes Sent
# group(8): Referrer URL
# group(9): User Agent String
Export & Deployment Actions1-click clipboard transfer, shareable URL hash, and local file downloads.

Built by Gordon Geraghty, Head of Performance MediaClient-Side Log Parser Engine
03 Deployment & Export

Run Turnkey Log Analysis Script

Download the pre-configured Python or Node.js script and execute against your server log files to output crawl budget analytics and spoofed bot reports.

Why Reverse DNS Verification Matters

Scrapers and rogue bots frequently forge Googlebot User-Agent strings. Validating crawler IPs against reverse DNS ensures that your log file analysis reflects real search engine crawl patterns rather than automated scrapers.

Implementation Code & Script

Googlebot Reverse DNS Verification Scriptverify-googlebot.mjsjavascript

Verifies whether a crawler IP belongs to genuine Googlebot using reverse and forward DNS lookups.

import dns from 'dns/promises';

export async function isGenuineGooglebot(ip) {
  try {
    const hostnames = await dns.reverse(ip);
    for (const host of hostnames) {
      if (host.endsWith('.googlebot.com') || host.endsWith('.google.com')) {
        const resolvedIps = await dns.resolve(host);
        if (resolvedIps.includes(ip)) {
          return true;
        }
      }
    }
    return false;
  } catch (err) {
    return false;
  }
}
04 QA & Verification Guide

Log Analysis & Bot Detection QA

Verify genuine search engine crawler IPs and detect scrapers spoofing Googlebot user-agent strings.

Pre-Production Verification Checklist

✓
Confirm fcRDNS on Googlebot User-Agents

Verify reverse DNS lookups resolve to *.googlebot.com and forward lookup matches origin IP.

✓
Identify 4xx / 5xx Waste in Crawl Budget

Confirm less than 2% of Googlebot requests return 404, 301 chains, or 500 server errors.

✓
Validate High-Frequency Crawl Directories

Verify priority product and category sections receive 80%+ of total search crawler volume.

Terminal Diagnostic & Debug Commands

Manual fcRDNS Verification (host CLI)bash

Verifies bidirectional DNS integrity to authenticate legitimate Googlebot IP addresses.

host 66.249.66.1 && host crawl-66-249-66-1.googlebot.com

Failure Remediation & Troubleshooting

Issue: High Volume of Spoofed Googlebot Crawls

Cause: Malicious scrapers copying Googlebot User-Agent to bypass Cloudflare rate limits.

Fix: Enable Cloudflare "Verified Bots" firewall rule to automatically drop spoofed bot requests at edge.

Issue: Googlebot Spending 50%+ Budget on Parameter URLs

Cause: Faceted search internal links are generating unique URLs for every filter click.

Fix: Disallow facet query parameters in robots.txt and implement URL parameter canonicalization.

How to cite and attribute this tool

MIT Licence

This resource is free, open and un-gated under the MIT Open Source Licence. You are encouraged to use, integrate and cite it with attribution:

Geraghty, G. (2026). Edge Server Log File Crawler & Bot Verification Script. Gordon Geraghty Resources Hub. https://gordongeraghty.com/resources/technical-seo/log-file-analysis-script
BibTeX Format
@misc{geraghty_log_file_analysis_script,
  author = {Geraghty, Gordon},
  title = {Edge Server Log File Crawler & Bot Verification Script},
  year = {2026},
  url = {https://gordongeraghty.com/resources/technical-seo/log-file-analysis-script},
  note = {Head of Performance Media, Empire Amplify}
}

Changelog & Version History

  • v1.0.0Initial release of edge log parser with reverse DNS bot verification.