Thumbnail for Raw data and analysis code for the study “Project 2025 as a technocratic blueprint: A corpus-based linguistic analysis of conservative governance discourse”

Raw data and analysis code for the study “Project 2025 as a technocratic blueprint: A corpus-based linguistic analysis of conservative governance discourse”

Important usage condition: Use and redistribution are governed by CC BY 4.0. Review and comply with the license before using the data.
Persistent identifier
doi:10.60507/FK2/BK14F8
Published version
1.0
Publication date
2025-11-11
License
CC BY 4.0

Description

This repository contains all data and scripts used for the study “Project 2025 as a Technocratic Blueprint: A Corpus-Based Linguistic Analysis of Conservative Governance Discourse” (Schilling & Fuchs, 2025). The study investigates the language of the Heritage Foundation’s Project 2025, a 900-page conservative policy blueprint, using methods from Corpus-Assisted Discourse Studies (CADS), Political Discourse Analysis (PDA), and psycholinguistic text analysis. The corpus includes Project 2025 and Democratic and Republican Party platforms (2016–2024). The dataset includes: Raw data (/data/raw_data/): full texts of Project 2025 and party platforms in CSV format. Processed data (/data/raw_data/Project2025_lemmaPOS.csv): tokenized, lemmatized, and POS-tagged text. Keyness results (/data/keyness/): unigram and bigram keyness calculations (log-likelihood, log-ratio). Collocation results (/data/collocations/): top 10 adjective, noun, and verb collocates per node. LIWC results (/data/liwc/): LIWC-22 category scores for each corpus. Analysis scripts (/code/): R Markdown file (analysis_project2025.Rmd) and two Python scripts for collocation and keyness analysis. All files are in UTF-8 plain-text format. The dataset contains no personal, sensitive, or proprietary data and derives entirely from publicly accessible political documents.

Creators

Keywords

Project 2025, Heritage Foundation, political discourse, policy communication, corpus linguistics, discourse analysis, collocation analysis, keyness, LIWC

Files

Before downloading: Use and redistribution are governed by CC BY 4.0. Review and comply with the license before using the data.
FileTypeBytesChecksum
keyness_2gram_pos_manifesto_rep_overuse.csvtext/csv13103889MD5 fc2cee9dd32ff05bc68acaabc9a260a7
collocation_analysis.pytext/x-python-script20533MD5 bb3e93a961f46a78fe49671abfec56ee
P2025_colloc_ADJ_w7_top10_per_node.csvtext/csv7007MD5 35f5564343e279ba2f9687dc354fd63e
keyness_1gram_pos_manifesto_dem_overuse.csvtext/csv1057511MD5 c96c120d14dadcb655181244fbc64de3
P2025_colloc_VERB_w7_top10_per_node.csvtext/csv7038MD5 fec75c16cb53fca40d7c5ed5f807e9c4
keyness_1gram_pos_manifesto_rep_overuse.csvtext/csv1058835MD5 206e1a07b39133218c461ccf7a4082f4
keyness_analysis.pytext/x-python-script9106MD5 fdba36fc4bbe3e9d0daf93a0cd6fc599
keyness_2gram_pos_manifesto_dem_overuse.csvtext/csv12827129MD5 220625d115ab7262e2764383921ff8f6
P2025_colloc_NOUN_w7_top10_per_node.csvtext/csv7110MD5 ac2033a4ca681c8e26ca0ab893bd955c
Platforms_Democrats_LIWC_subset.csvtext/csv916161MD5 1a33450f5ce0ac9dd0e34d60ed0c350f
Project2025_LIWC_subset.csvtext/csv2310611MD5 75c059d442c8a215e95d4dded4b186c8
Platforms_Republicans_LIWC_subset.csvtext/csv330359MD5 641ccc8365b1afd101375b50a1a86a3a
analysis_project2025.Rmdtext/x-r-notebook26326MD5 136d84af88837928bbedb7225e6bf0bd
Project2025.csvtext/csv2154423MD5 4662cb4b2b1301d0d40c83d077959ea7
Platforms_Democrats.csvtext/csv846946MD5 aabcf833be06563a10bc99671e8f02f8
Platforms_Republicans.csvtext/csv303972MD5 a18aca13011427bccdb4c70b35e8608d
Project2025_lemmaPOS.csvtext/csv4200254MD5 24a16b6e1b93862c3a3975bca7d55e5d
README_Project2025.mdtext/markdown9188MD5 2e5e8f4560c10cd97b8f04adcbe4e023

Citation

Schilling, Julia; Fuchs, Robert, 2025-11-11, Raw data and analysis code for the study “Project 2025 as a technocratic blueprint: A corpus-based linguistic analysis of conservative governance discourse”, doi:10.60507/FK2/BK14F8, V1.0

Additional Dataverse fields

Export metadata

Static metadata exports available for this published dataset version:

Complete Dataverse metadata

Expected crawler behaviour

Use a stable, truthful User-Agent with product/version and a working contact URL. Across all IP addresses and HTTP connections used by one crawler identity, allow no more than 5 requests in flight and wait at least 20 seconds between request starts. Crawl URLs listed in the catalog sitemap, including file pages and download URLs when they are published, use conditional requests, honor Retry-After, and apply exponential backoff after errors.

The welcome page may link to the interactive repository for human navigation. Automated clients must not treat that human link as a catalog crawl target.

Read the live machine-readable crawler policy before and during a crawl. Stop crawling when it reports CPU or memory utilization at or above 80% and 80% respectively.