API¶
django-crawl can also be called from Python with django_crawl.crawl().
This is most useful for crawling your site within your test suite, where you can use your factories or fixtures to populate the database, and broken pages fail the test with their exceptions.
Minimal usage looks like:
from django.test import TestCase
import django_crawl
class CrawlTests(TestCase):
def test_crawl(self):
django_crawl.crawl()
- django_crawl.crawl(*start_urls, client=None, depth=5, max_urls=1000, max_query_variants=10, exclude=(), on_response=None, check=True)[source]¶
Crawl the site in-process and return a
CrawlResult, like the crawl management command.- Parameters:
start_urls – URL paths to start crawling from. Defaults to
/. Invalid start URLs raiseValueError.client – The Django test client instance to crawl with. Defaults to a fresh
django_crawl.CrawlClient, a testClientsubclass that also serves static files, likerunserverdoes. Pass in a client to customize log in, request headers, or other behaviour, per the below example.depth – Maximum link depth to crawl.
0means crawl only the start URLs without following links.max_urls – Maximum number of URLs to request.
max_query_variants – Maximum number of query string variants to crawl per path, or
Nonefor unlimited.exclude – Regular expressions, as strings or compiled patterns. Discovered URLs matching any pattern are skipped; start URLs are always crawled.
on_response – A callable invoked with each response, for making custom checks. Exceptions it raises, such as from
assertstatements, are recorded as errors for that URL.check – Whether to raise an
ExceptionGroupat the end if any errors were found (the default). PassFalseto instead return the errors in the result.
Unlike the management command,
crawl():does not log in automatically.
produces no output.
stops at the first
ValueErrorfor an invalid start URL rather than reporting errors as it goes.
Logging in¶
To crawl pages that require authentication, create a client, log it in, and pass it in:
from django.contrib.auth.models import User
from django.test import TestCase
import django_crawl
class CrawlTests(TestCase):
@classmethod
def setUpTestData(cls):
cls.admin = User.objects.create_superuser(username="admin")
def test_crawl(self):
client = django_crawl.CrawlClient()
client.force_login(self.admin)
django_crawl.crawl("/", "/admin/", client=client)
The client can also customize requests in other ways, such as sending extra headers:
client = django_crawl.CrawlClient(headers={"accept-language": "de"})
django_crawl.crawl(client=client)
Checking responses¶
Use on_response to check every response, like the management command’s -c option:
class CrawlTests(TestCase):
def test_crawl(self):
def check_csp(response):
assert (
"content-security-policy" in response.headers
), response.wsgi_request.path
django_crawl.crawl(on_response=check_csp)
Results¶
With check=True (the default), broken pages raise an ExceptionGroup.
Each contained exception is the original exception the page raised, with the URL attached as a note, or a django_crawl.ResponseError for plain HTTP error statuses, like HTTP 404 Not Found: /dead-link/.
(On Python 3.10, original exceptions are wrapped in ResponseError with __cause__ set, since exception notes require Python 3.11.)
With check=False, crawl() returns a CrawlResult without raising:
result = django_crawl.crawl(check=False)
CrawlResult has three attributes:
count: the number of URLs crawled.errors: a list ofCrawlErrorinstances, each withurl,message, andexc_infoattributes.stop_reason: adjango_crawl.StopReasonenum value,NO_MORE_LINKSorMAX_URLS.