Raw data and analysis code for the study “Project 2025 as a technocratic blueprint: A corpus-based linguistic analysis of conservative governance discourse”
Open this dataset in the live repository
- Persistent identifier
- doi:10.60507/FK2/BK14F8
- Published version
- 1.0
- Publication date
- 2025-11-11
- License
- CC BY 4.0
Description
This repository contains all data and scripts used for the study “Project 2025 as a Technocratic Blueprint: A Corpus-Based Linguistic Analysis of Conservative Governance Discourse” (Schilling & Fuchs, 2025). The study investigates the language of the Heritage Foundation’s Project 2025, a 900-page conservative policy blueprint, using methods from Corpus-Assisted Discourse Studies (CADS), Political Discourse Analysis (PDA), and psycholinguistic text analysis. The corpus includes Project 2025 and Democratic and Republican Party platforms (2016–2024). The dataset includes: Raw data (/data/raw_data/): full texts of Project 2025 and party platforms in CSV format. Processed data (/data/raw_data/Project2025_lemmaPOS.csv): tokenized, lemmatized, and POS-tagged text. Keyness results (/data/keyness/): unigram and bigram keyness calculations (log-likelihood, log-ratio). Collocation results (/data/collocations/): top 10 adjective, noun, and verb collocates per node. LIWC results (/data/liwc/): LIWC-22 category scores for each corpus. Analysis scripts (/code/): R Markdown file (analysis_project2025.Rmd) and two Python scripts for collocation and keyness analysis. All files are in UTF-8 plain-text format. The dataset contains no personal, sensitive, or proprietary data and derives entirely from publicly accessible political documents.
Creators
- Schilling, Julia
- Fuchs, Robert
Keywords
Project 2025, Heritage Foundation, political discourse, policy communication, corpus linguistics, discourse analysis, collocation analysis, keyness, LIWC
Files
| File | Type | Bytes | Checksum |
|---|---|---|---|
| keyness_2gram_pos_manifesto_rep_overuse.csv | text/csv | 13103889 | MD5 fc2cee9dd32ff05bc68acaabc9a260a7 |
| collocation_analysis.py | text/x-python-script | 20533 | MD5 bb3e93a961f46a78fe49671abfec56ee |
| P2025_colloc_ADJ_w7_top10_per_node.csv | text/csv | 7007 | MD5 35f5564343e279ba2f9687dc354fd63e |
| keyness_1gram_pos_manifesto_dem_overuse.csv | text/csv | 1057511 | MD5 c96c120d14dadcb655181244fbc64de3 |
| P2025_colloc_VERB_w7_top10_per_node.csv | text/csv | 7038 | MD5 fec75c16cb53fca40d7c5ed5f807e9c4 |
| keyness_1gram_pos_manifesto_rep_overuse.csv | text/csv | 1058835 | MD5 206e1a07b39133218c461ccf7a4082f4 |
| keyness_analysis.py | text/x-python-script | 9106 | MD5 fdba36fc4bbe3e9d0daf93a0cd6fc599 |
| keyness_2gram_pos_manifesto_dem_overuse.csv | text/csv | 12827129 | MD5 220625d115ab7262e2764383921ff8f6 |
| P2025_colloc_NOUN_w7_top10_per_node.csv | text/csv | 7110 | MD5 ac2033a4ca681c8e26ca0ab893bd955c |
| Platforms_Democrats_LIWC_subset.csv | text/csv | 916161 | MD5 1a33450f5ce0ac9dd0e34d60ed0c350f |
| Project2025_LIWC_subset.csv | text/csv | 2310611 | MD5 75c059d442c8a215e95d4dded4b186c8 |
| Platforms_Republicans_LIWC_subset.csv | text/csv | 330359 | MD5 641ccc8365b1afd101375b50a1a86a3a |
| analysis_project2025.Rmd | text/x-r-notebook | 26326 | MD5 136d84af88837928bbedb7225e6bf0bd |
| Project2025.csv | text/csv | 2154423 | MD5 4662cb4b2b1301d0d40c83d077959ea7 |
| Platforms_Democrats.csv | text/csv | 846946 | MD5 aabcf833be06563a10bc99671e8f02f8 |
| Platforms_Republicans.csv | text/csv | 303972 | MD5 a18aca13011427bccdb4c70b35e8608d |
| Project2025_lemmaPOS.csv | text/csv | 4200254 | MD5 24a16b6e1b93862c3a3975bca7d55e5d |
| README_Project2025.md | text/markdown | 9188 | MD5 2e5e8f4560c10cd97b8f04adcbe4e023 |
Citation
Schilling, Julia; Fuchs, Robert, 2025-11-11, Raw data and analysis code for the study “Project 2025 as a technocratic blueprint: A corpus-based linguistic analysis of conservative governance discourse”, doi:10.60507/FK2/BK14F8, V1.0
Additional Dataverse fields
| Id | 517 |
|---|---|
| Dataset Type | dataset |
| Internal Version Number | 31 |
| Latest Version Publishing State | RELEASED |
| Release Time | 2025-11-11T12:47:28Z |
| Create Time | 2025-11-05T19:06:46Z |
| Citation Date | 2025-11-11 |
| File Access Request | True |
Export metadata
Static metadata exports available for this published dataset version:
Complete Dataverse metadata
Expected crawler behaviour
Use a stable, truthful User-Agent with product/version and a working contact URL. Across all IP addresses and HTTP connections used by one crawler identity, allow no more than 5 requests in flight and wait at least 20 seconds between request starts. Crawl URLs listed in the catalog sitemap, including file pages and download URLs when they are published, use conditional requests, honor Retry-After, and apply exponential backoff after errors.
The welcome page may link to the interactive repository for human navigation. Automated clients must not treat that human link as a catalog crawl target.
Read the live machine-readable crawler policy before and during a crawl. Stop crawling when it reports CPU or memory utilization at or above 80% and 80% respectively.
