Skip to main content
Ctrl+K
Trafilatura 2.2.0 documentation - Home Trafilatura 2.2.0 documentation - Home
  • Installation
  • Quickstart
  • Usage
  • Tutorials
  • FAQ
    • Troubleshooting
    • How extraction works
    • Benchmarks and evaluation
    • Core functions
    • Settings and customization
    • Deprecations and migration
    • Uses & citations
    • Running the tests
    • Background
    • Blog
  • GitHub
  • Installation
  • Quickstart
  • Usage
  • Tutorials
  • FAQ
  • Troubleshooting
  • How extraction works
  • Benchmarks and evaluation
  • Core functions
  • Settings and customization
  • Deprecations and migration
  • Uses & citations
  • Running the tests
  • Background
  • Blog
  • GitHub

Section Navigation

  • Compendium: Web texts in linguistics and humanities
  • Finding sources for web corpora
  • Working with corpus data
  • Background

Background#

The pages below provide background information on scientific approaches to web data collection and processing, corpus linguistics, digital humanities, and natural language processing.

  • Compendium: Web texts in linguistics and humanities
    • Web corpora as scientific objects
    • Corpus types and resulting methods
    • Corpus construction steps
    • Methodological issues
    • References
  • Finding sources for web corpora
    • From link lists to web corpora
    • Existing resources
    • Search engines
    • Selecting random documents from the Web
    • Social networks
    • Remarks
    • References
  • Working with corpus data
    • Generic solutions in Python
    • Formats and software used in corpus linguistics
    • Generic NLP solutions

previous

Running the tests

next

Compendium: Web texts in linguistics and humanities

Show Source

© Copyright 2025, Adrien Barbaresi.

Built with the PyData Sphinx Theme 0.20.0.