Tearing it up with Django.
Tracking the progress of the site rebuild in Django and its ability to handle heavy bot traffic.
Crawlers Gonna Crawl
The final redesign of this site is complete. For today, anyways.
I've been watching the logs of this site and my other WordPress site. I made an interesting discovery:
Invasive bots DO NOT stick around Django-based sites (which this is), but rather authentic exploratory bots tend to trickle in and do their thing.
Django, being a very efficient system, handles them without any real problem.
WordPress -- The Invasive Bot Magnet
WORDPRESS, however, is another story. Once a site is identified as WordPress, the hacker bots move right in and crawl it constantly, testing all the common exposed .php files, running various weakness tests, etc.
This necessitates the installation of protection plugins, which of course use more memory. The more memory that is used, the more you have to cache -- another plugin. The effect compounds: Larger WordPress site, more attention, more stress on the server. More protection.
I remembered very quickly why I DO NOT like WordPress. It's an expendable site, however.
Bots Found on This Site (Django)
Python-urllib/3.14[Neutral] — Generic Python HTTP client used by scripts and automation; not a specific identifiable bot.AhrefsBot/7.0[Invasive] — Ahrefs crawler for backlink indexing, SEO analysis, and web-link collection.Amazonbot/0.1[Invasive] — Amazon crawler used for Amazon services and potentially AI-related data collection.PerplexityBot/1.0[Beneficial] — Perplexity crawler used to discover and index web content for search and answer systems.GoogleOther[Neutral] — Google's general-purpose crawler for internal products and non-Search-specific crawling.MJ12bot/v1.4.8[Invasive] — Majestic crawler used to build its backlink and web-link index.PetalBot[Beneficial] — Huawei crawler associated with Petal Search and web indexing.Googlebot/2.1[Beneficial] — Google's primary crawler for discovering and indexing pages in Google Search.Bytespider[Invasive] — ByteDance crawler used for web indexing and large-scale data collection.ClaudeBot/1.0[Invasive] — Anthropic crawler that collects public web content that may be used for Claude model development and training.Python-urllib/3.12[Neutral] — Generic Python HTTP client used by scripts and automation; not a specific identifiable bot.SERankingBacklinksBot/1.0[Invasive] — SE Ranking crawler that collects links for backlink analysis.facebookexternalhit/1.1[Beneficial] — Meta/Facebook fetcher used to generate link previews when URLs are shared.Sogou[Beneficial] — Crawler for the Sogou search engine.SemrushBot/7~bl[Invasive] — Semrush crawler used primarily for backlink and SEO data collection.Python-urllib/3.13[Neutral] — Generic Python HTTP client used by scripts and automation; not a specific identifiable bot.OAI-SearchBot/1.4[Beneficial] — OpenAI crawler used to discover and index pages for ChatGPT Search.Barkrowler/0.9[Invasive] — Babbar crawler used to collect links and page metadata for SEO/link analysis.bingbot/2.0[Beneficial] — Microsoft's crawler for discovering and indexing content in Bing Search.Baiduspider/2.0[Beneficial] — Baidu's primary web crawler for its search engine index.SiteAuditBot/0.97[Invasive] — Semrush crawler used for technical SEO and site-audit scans.DuckAssistBot/1.2[Beneficial] — DuckDuckGo crawler that retrieves web content for AI-assisted search answers.ExaSearchBot/1.0[Beneficial] — Exa crawler used to build its search and retrieval index.ChatGPT-User/1.0[Beneficial] — OpenAI user-triggered fetcher used when ChatGPT accesses a page on behalf of a user.GPTBot/1.4[Invasive] — OpenAI crawler that collects public web content that may be used to train generative AI models.DuckDuckBot/1.1[Beneficial] — DuckDuckGo crawler used to build and refresh its search index.YouBot/1.0[Beneficial] — You.com crawler used for search indexing and AI-powered search results.DuckDuckBot/1.0[Beneficial] — Older DuckDuckGo crawler used for search indexing.MistralAI-User/1.0[Beneficial] — Mistral user-triggered fetcher that retrieves pages when a user asks its AI to access web information.cohere-ai[Invasive] — Legacy or claimed Cohere crawler identity associated with AI-oriented web collection.Google-Extended[Invasive] — Google control token associated with Gemini training and grounding rather than ordinary Google Search indexing.Baiduspider-render/2.0[Beneficial] — Baidu rendering crawler used to process JavaScript-heavy pages for search.CCBot/2.0[Invasive] — Common Crawl crawler used to build large public web datasets.Kimi-SearchBot/1.0[Beneficial] — Search crawler associated with Moonshot AI's Kimi search and retrieval services.GrokBot/1.0[Invasive] — Claimed xAI/Grok crawler associated with AI web retrieval or data collection.ChatGLM-Spider/1.0[Invasive] — AI crawler associated with the ChatGLM/Zhipu ecosystem and likely used for large-scale web collection.MoonshotBot/1.0[Invasive] — Crawler associated with Moonshot AI/Kimi for AI-related web collection.PanguBot/1.0[Invasive] — Undocumented crawler using Huawei Pangu branding and apparently associated with AI-oriented collection.Qwenbot/1.0[Invasive] — Crawler associated with Alibaba's Qwen AI ecosystem and likely used for web retrieval or AI data collection.xAI-Grok/1.0[Invasive] — Claimed xAI/Grok crawler associated with AI retrieval or data collection.LohiSoftBot/1.0[Beneficial] — LohiSoft crawler used to build a search index from page metadata, headings, URLs, and links.Googlebot-Image/1.0[Beneficial] — Google crawler dedicated to discovering and indexing images for Google Images.Claude-User/1.0[Beneficial] — Anthropic user-triggered fetcher that accesses pages when a Claude user requests web information.DeepSeekBot/1.0[Invasive] — Claimed DeepSeek AI crawler associated with automated web collection.Perplexity-User/1.0[Beneficial] — Perplexity user-triggered fetcher that retrieves pages while answering a user's query.YiBot/1.0[Invasive] — AI crawler associated with the Yi/01.AI ecosystem.CensysInspect/1.1[Neutral] — Censys internet scanner used to discover and measure publicly exposed servers and services.KimiBot/1.0[Invasive] — General crawler associated with Moonshot AI's Kimi services.KeenableBot/1.0[Invasive] — Keenable.ai crawler used to collect or index public web content for AI/search systems.Hunyuan/1.0[Invasive] — Claimed crawler associated with Tencent's Hunyuan AI ecosystem.DuckAssistBot/1.1[Beneficial] — Older DuckDuckGo crawler used to retrieve web content for AI-assisted search answers.Meta-ExternalAgent/1.0[Invasive] — Meta crawler used for AI model training and direct content indexing for Meta products.OAI-SearchBot/1.0[Beneficial] — Older OpenAI crawler used to discover and index pages for ChatGPT Search.Synapse[Neutral] — Matrix Synapse HTTP client commonly used for URL previews, remote media, or other Matrix-related fetching.coccocbot-image/1.0[Beneficial] — Cốc Cốc crawler used to discover and index images for its search engine.SemrushBot-BA[Invasive] — Semrush crawler specifically used for backlink-audit data.NostrichBot/1.0[Neutral] — Nostr-related automated crawler; exact operator and purpose are unclear.coccocbot-web/1.0[Beneficial] — Cốc Cốc crawler used to discover and index normal web pages for search.CFNetwork/1568.200.51[Neutral] — Apple's CFNetwork HTTP stack used by macOS/iOS applications; not a specific bot identity.DotBot/1.2[Invasive] — Moz crawler used to build link and SEO datasets.Qwantbot/1.0[Beneficial] — Qwant crawler used to discover and index pages for the Qwant search engine.InternetMeasurement/1.0[Neutral] — Internet scanner used to discover and measure publicly exposed network services.YandexBot/3.0[Beneficial] — Yandex crawler used to discover and index pages for Yandex Search.PubkyWebIndex/0.1[Beneficial] — Crawler associated with Pubky web indexing.PipericBot/1.0[Invasive] — Piperic Business Intelligence crawler used to index public company and website information.Python/3.12[Neutral] — Generic Python application user-agent; usually a custom script, scraper, monitor, or automation.Google-PageRenderer[Beneficial] — Google rendering service used to load and render pages, particularly JavaScript-dependent content.YandexFavicons/1.0[Beneficial] — Yandex crawler dedicated to retrieving website favicons.WhatsApp/2[Beneficial] — WhatsApp fetcher used primarily to retrieve URLs and generate link previews.BertinaAccessibilityProbe/0.1[Neutral] — Automated accessibility-testing probe.ListSignalBot/1.0[Invasive] — Automated crawler apparently used for business or lead-data discovery.CompanyMapBot/0.1[Invasive] — Automated crawler apparently used for company/business discovery and mapping.coinos-link-preview/1.0[Beneficial] — Coinos fetcher used to retrieve URLs and generate link previews.MJ12bot/v2.0.7[Invasive] — Newer Majestic crawler used to build and update its backlink and web-link index.wpbot/1.4[Neutral] — Unidentified WordPress-related bot whose purpose cannot reliably be determined from its user-agent.WordPress/6.2[Neutral] — WordPress HTTP client used for server-to-server requests such as APIs, feeds, pingbacks, updates, or scheduled tasks.GoogleImageProxy[Beneficial] — Google image proxy/cache used to retrieve and serve remote images through Google services.CFNetwork/3860.700.2[Neutral] — Apple's CFNetwork HTTP stack used by macOS/iOS applications; not a specific bot identity.Google/1.0[Neutral] — Generic Google-labelled user-agent whose exact service cannot be determined from the string alone.Python/3.10[Neutral] — Generic Python application user-agent, normally a custom script, scraper, monitor, or automation.bot/1.1[Neutral] — Completely generic bot identifier; operator and purpose cannot be determined.TheMoneyWatch/1.0[Neutral] — Undocumented crawler identifying itself as TheMoneyWatch, apparently related to finance monitoring or indexing.
Bots Found on Other Site (WordPress)
KeenableBot/1.0[Invasive] — AI/search crawler operated by Keenable; performs large-scale automated collection and indexing of public web content.Googlebot/2.1[Beneficial] — Google's primary crawler for discovering and indexing pages in Google Search.GoogleOther[Neutral] — Google crawler used for non-standard Google products and internal crawling outside normal Google Search indexing.MJ12bot/v1.4.8[Invasive] — Majestic crawler used to build its commercial backlink and web-link database.WordPress/7.1.2[Neutral] — WordPress HTTP client used for automated server-to-server requests such as cron, APIs, feeds, pingbacks, plugin activity, and loopback requests.SERankingBacklinksBot/1.0[Invasive] — SE Ranking crawler used to collect backlink data for commercial SEO analysis.Amazonbot/0.1[Invasive] — Amazon crawler used for Amazon services and AI-related web collection.DataForSeoBot/1.0[Invasive] — DataForSEO crawler used to build commercial backlink and SEO datasets.SemrushBot/7~bl[Invasive] — Semrush crawler used to collect backlink and SEO intelligence.AwarioBot/1.0[Invasive] — Awario crawler used for social listening, brand monitoring, sentiment analysis, and lead discovery.PetalBot[Beneficial] — Huawei crawler used to index pages for Petal Search.AhrefsBot/7.0[Invasive] — Ahrefs crawler used to build commercial backlink and SEO datasets.DotBot/1.2[Invasive] — Moz crawler used to collect link and SEO data.WordPress.com[Neutral] — WordPress.com automated client used by WordPress-hosted and Jetpack-related services.ExaSearchBot/1.0[Beneficial] — Exa crawler used to build a search and retrieval index consumed by search and AI applications.PubkyWebIndex/0.1[Beneficial] — Pubky crawler used to index publicly accessible web and Pubky content.WordPress/6.3[Neutral] — WordPress HTTP client performing automated server-to-server requests.WordPress/6.4[Neutral] — WordPress HTTP client performing automated server-to-server requests.WordPress/6.1[Neutral] — WordPress HTTP client performing automated server-to-server requests.WordPress/6.2[Neutral] — WordPress HTTP client performing automated server-to-server requests.bingbot/2.0[Beneficial] — Microsoft's crawler used to discover and index pages for Bing Search.meta-externalagent/1.1[Invasive] — Meta crawler used for AI-related data collection and indexing for Meta products.facebookexternalhit/1.1[Beneficial] — Facebook/Meta fetcher used to retrieve pages and generate link previews when URLs are shared.OAI-SearchBot/1.4[Beneficial] — OpenAI crawler used to discover and index web pages for ChatGPT Search.CFNetwork/1568.200.51[Neutral] — Apple's generic CFNetwork HTTP client; may represent an application, preview fetch, or automated request rather than a specific crawler.Baiduspider-render/2.0[Beneficial] — Baidu rendering crawler used to process JavaScript-dependent pages for search indexing.CFNetwork/3860.700.2[Neutral] — Apple's generic CFNetwork HTTP client; exact application or purpose cannot be identified from the user-agent alone.ClaudeBot/1.0[Invasive] — Anthropic crawler used to collect public web content that may be used for Claude development and training.GPTBot/1.4[Invasive] — OpenAI crawler used to collect public web content that may be used for generative-AI model development.CFNetwork/3896.100.1.2.1[Neutral] — Apple's generic CFNetwork HTTP client; exact application or purpose is unknown.CFNetwork/3860.600.12[Neutral] — Apple's generic CFNetwork HTTP client; exact application or purpose is unknown.CFNetwork/3860.700.1[Neutral] — Apple's generic CFNetwork HTTP client; exact application or purpose is unknown.CFNetwork/3860.300.31[Neutral] — Apple's generic CFNetwork HTTP client; exact application or purpose is unknown.PerplexityBot/1.0[Beneficial] — Perplexity crawler used to index web content for its search and answer systems.AdsBot-Google[Beneficial] — Google crawler used to inspect landing pages for Google Ads quality and policy purposes.QlyzeBot/1.0[Invasive] — Automated data crawler; exact operator documentation is limited and it provides no obvious direct search-discovery benefit.ChatGPT-User/1.0[Beneficial] — OpenAI fetcher activated when a ChatGPT user requests access to a particular webpage.l9scan/2.0[Invasive] — Automated internet scanner used to probe websites and exposed services rather than index them for users.NostrichBot/1.0[Neutral] — Nostr-related automated crawler; exact indexing purpose and operator are unclear.CFNetwork/1410.1[Neutral] — Apple's generic CFNetwork HTTP client.CFNetwork/3826.600.41[Neutral] — Apple's generic CFNetwork HTTP client.CFNetwork/3826.600.41.2.1[Neutral] — Apple's generic CFNetwork HTTP client.LinkPreview/1.6[Beneficial] — Fetcher used to retrieve page metadata, images, and titles for link previews.client[Neutral] — Generic user-agent providing no reliable information about the software or purpose behind the request.Bytespider[Invasive] — ByteDance crawler used for large-scale web collection and services associated with ByteDance AI and search products.2140SocialBot/1.0[Neutral] — Social-oriented crawler or fetcher; exact operator and purpose are not sufficiently documented.facebookexternalhit[Beneficial] — Facebook/Meta URL fetcher used primarily to generate link previews.DuckAssistBot/1.2[Beneficial] — DuckDuckGo crawler used to obtain web information for AI-assisted answers.WhatsApp/2[Beneficial] — WhatsApp fetcher used to retrieve URLs and create link previews when pages are shared.ImageFetcher/9.0[Neutral] — Generic image-retrieval client; exact operator and reason for fetching the images are unclear.AIWebIndex/2.0[Beneficial] — Open AI-readable web indexing protocol used by compliant AI systems to retrieve structured representations of public webpages.CFNetwork/1568.100.1.2.1[Neutral] — Apple's generic CFNetwork HTTP client.IronFountain-Leads/1.0[Invasive] — Automated lead-generation crawler used to collect business or contact information.Googlebot-Image/1.0[Beneficial] — Google crawler used to discover and index images for Google Images.AdsBot-Google-Mobile[Beneficial] — Google's mobile landing-page crawler used by Google Ads for quality and policy evaluation.DuckDuckBot/1.0[Beneficial] — DuckDuckGo crawler used to discover and index webpages for search.OAI-SearchBot/1.0[Beneficial] — Older OpenAI crawler used to index webpages for ChatGPT Search.DuckDuckBot/1.1[Beneficial] — DuckDuckGo crawler used to discover and index webpages for search.CFNetwork/3860.500.112[Neutral] — Apple's generic CFNetwork HTTP client.CFNetwork/3860.600.21[Neutral] — Apple's generic CFNetwork HTTP client.CensysInspect/1.1[Neutral] — Censys scanner used to identify and measure publicly exposed internet servers and services.Keenable-User/1.0[Beneficial] — Keenable fetcher apparently activated on behalf of an end user rather than performing autonomous bulk crawling.CFNetwork/3860.200.71[Neutral] — Apple's generic CFNetwork HTTP client.Baiduspider/2.0[Beneficial] — Baidu crawler used to discover and index webpages for Baidu Search.CFNetwork/3860.400.51[Neutral] — Apple's generic CFNetwork HTTP client.SemrushBot-BA[Invasive] — Semrush crawler used to gather backlink information for its Backlink Audit products.LiberMedia/2.0[Neutral] — Automated media-related fetcher whose exact operator and collection purpose are unclear.ProspectDB/1.0[Invasive] — Automated prospecting crawler apparently used to collect business and lead information.ZapCookingBot/1.0[Neutral] — Automated crawler whose exact operator and purpose are insufficiently documented.CFNetwork/3826.500.131[Neutral] — Apple's generic CFNetwork HTTP client.CFNetwork/1498.700.2[Neutral] — Apple's generic CFNetwork HTTP client.Nostrich/1.0[Neutral] — Nostr-related client or fetcher; exact purpose cannot be determined from the user-agent alone.CFNetwork/3860.100.1[Neutral] — Apple's generic CFNetwork HTTP client.DataSeekerBot/1.0[Invasive] — Automated crawler apparently designed to discover and collect web data rather than provide direct search traffic.netEstate[Invasive] — netEstate crawler used to collect publicly listed contact and company information from websites.CFNetwork/3890.100.1[Neutral] — Apple's generic CFNetwork HTTP client.CompanyMapBot/0.1[Invasive] — Automated crawler apparently used to collect and map company or business information.python-httpx/0.28.1[Neutral] — Generic Python HTTP client commonly used by custom scripts, APIs, monitors, and scrapers.CFNetwork/3896.200.41[Neutral] — Apple's generic CFNetwork HTTP client.Querit-SearchBot/1.0[Beneficial] — Search-oriented crawler apparently used to build or query a web search index.TinEye-Web/2.0.1[Beneficial] — TinEye crawler used to discover images for reverse-image search.Python/3.11[Neutral] — Generic Python application user-agent, usually representing custom scripts or automation.CFNetwork/3826.400.120[Neutral] — Apple's generic CFNetwork HTTP client.probe[Invasive] — Generic probe identifier strongly suggesting automated testing or reconnaissance rather than normal indexing.Imwald/1.0[Neutral] — Undocumented automated client whose operator and purpose cannot reliably be identified.MJ12bot/v2.0.7[Invasive] — Newer Majestic crawler used to build and update its commercial backlink database.TwitterAndroid/12.30.0-prod.01[Beneficial] — X/Twitter Android application request, potentially associated with a user opening or sharing a link rather than autonomous crawling.payouts[Neutral] — Generic automated client identifier; its operator and purpose cannot be determined from the user-agent.meta-webindexer/1.1[Invasive] — Meta web-indexing crawler used to collect public webpage information for Meta services.bot/1.1[Neutral] — Completely generic bot identifier with no reliable operator or purpose information.research[Neutral] — Generic automated user-agent claiming a research purpose; operator and actual activity are unknown.ListSignalBot/1.0[Invasive] — Automated crawler apparently used for business-data or lead-list discovery.Tool[Neutral] — Generic automated client identifier with no identifiable operator or purpose.Google/569[Neutral] — Generic Google application user-agent; the exact Google service responsible cannot be determined from this string.CFNetwork/1402.0.8[Neutral] — Apple's generic CFNetwork HTTP client.ObeliskBot/1.0[Neutral] — Automated crawler whose operator and purpose are not sufficiently documented to classify more specifically.Crawler/0.2[Neutral] — Generic crawler identifier with no reliable operator or purpose information.RegieBrainBot/0.1[Invasive] — Automated crawler associated with sales/prospecting intelligence and lead-generation workflows.K7MLWCBot/1.0[Neutral] — Undocumented crawler whose operator and purpose cannot reliably be determined.CFNetwork/1335.0.3.4[Neutral] — Apple's generic CFNetwork HTTP client.
The Comparison:
| Django | WordPress | WP vs Django | |
|---|---|---|---|
| Invasive requests | 10,420 | 194,589 | 18.7× more |
| Invasive bandwidth | ~149.8 MiB | ~7.39 GiB | 50.5× more |
| Invasive share of crawler requests | 29.5% | 64.8% | +35.3 percentage points |
So WordPress wins the invasive bot contest by a landslide. And guess what? The WordPress site is only a month old --- this site is almost a year old.
This should tell you that running WordPress makes you a target.
The Conclusion
Don't use WordPress if you can avoid it. Django is fine-tuned, rock solid, secure. WordPress is clunky, inefficient, and generally not secure. (They say it is, but after plugins, it is not).
As a post-note -- returning to WordPress in 2026 showed me just how badly the plugin community has degraded. Everything is a polished upsell with minimally available good features.
Needless to say, I've got another WordPress -> Django conversion coming up soon.
- Categories: Web Development
- Tags: #Django, #WordPress, #Site Development, #Bot Traffic