Ask crawlers to leave in robots.txt first; GPTBot, CCBot and AhrefsBot say they obey it. This file handles the rest, plus scanners and image thieves:
# Needs: mod_setenvif, mod_rewrite; AllowOverride AuthConfig FileInfo; Options FollowSymLinks
BrowserMatchNoCase "(AhrefsBot|SemrushBot|MJ12bot|DotBot|PetalBot|Bytespider)" bad_bot
BrowserMatchNoCase "(GPTBot|CCBot|sqlmap|nikto|masscan|zgrab)" bad_bot
BrowserMatch "^$" bad_bot
<RequireAll>
Require all granted
Require not env bad_bot
</RequireAll>
RewriteEngine On
RewriteRule (^|/)(wp-login|xmlrpc|phpinfo)\.php$ - [F]
RewriteCond %{HTTP_REFERER} !^$
RewriteCond %{HTTP_REFERER} !^https?://([^/]+\.)?example\.com(:\d+)?(/|$) [NC]
RewriteCond %{HTTP_REFERER} !^https?://(\w+\.)?(google|bing|duckduckgo)(\.com?)?(\.\w\w)?(/|$)
RewriteRule \.(avif|gif|jpe?g|png|webp)$ /img/hotlinked.svg [NC,L]Line 4 tags a request with no User-Agent, and lines 5 to 8 refuse tagged clients. Line 10 refuses WordPress 48 probes on a site without it. Lines 11 to 14 are Hotlink Protection's test, serving a placeholder instead of a 403:
AhrefsBot/7.0, / 403 text/html no User-Agent, / 403 text/html Firefox, /wp-login.php 403 text/html Firefox, /img/cat.jpg, from www.google.com 200 image/jpeg Firefox, /img/cat.jpg, from forum.example.net 200 image/svg+xml
A User-Agent is easily faked, so this stops only bots that name themselves.