Replication data for "Analyzing the Potency of Pretrained Transformer Models for Automated Program Repair"
Open this dataset in the live repository
- Persistent identifier
- doi:10.60507/FK2/O54LWP
- Published version
- 2.0
- Publication date
- 2024-07-15
- License
- CC BY 4.0
Description
This repository contains replication data for the paper "Analyzing the Potency of Pretrained Transformer Models for Automated Program Repair".<br> The dataset contains commits that indicate a bug fix from 200 repositories on Github. Scripts for fetching data associated with the commits are available (see Github repository linked under "Related Material"). <hr> File descriptions: <ul> <li><code>repositories.csv</code>: List of repository names with metadata</li> <li><code>commits.csv</code>: List of commit identifiers of the 200 repositories with most bug fixing commits</li> <li><code>commits_high_watch_count.csv</code>: List of commit identifiers of the 200 repositories with most bug fixing commits with a watch count of at least 50</li> </ul> <hr> Icon licensed under CC-BY by xinh.studio
Creators
- Leiwig, Maximilian
- Swierzy, Ben
- Bungartz, Christian
- Meier, Michael
Keywords
Program Repair, Machine Learning, Large Language Models
Files
| File | Type | Bytes | Checksum |
|---|---|---|---|
| repositories.csv | text/csv | 3144907 | MD5 f578867bd4854a1b228f335aa1b75d26 |
| commits_high_watch_count.csv | text/csv | 1567538 | MD5 64524b826cd44b2abf7064b794dbe6fe |
| commits.csv | text/csv | 6963520 | MD5 32a7d4bbb62765d708713b2ca8d6f5a8 |
| README.md | text/markdown | 1945 | MD5 bd4306963175ff314ccc46d7601a2b35 |
Citation
Leiwig, Maximilian; Swierzy, Ben; Bungartz, Christian; Meier, Michael, 2024-07-15, Replication data for "Analyzing the Potency of Pretrained Transformer Models for Automated Program Repair", doi:10.60507/FK2/O54LWP, V2.0
Additional Dataverse fields
| Id | 368 |
|---|---|
| Dataset Type | dataset |
| Internal Version Number | 7 |
| Latest Version Publishing State | RELEASED |
| Deaccession Link | Not supplied |
| Release Time | 2025-01-14T11:07:50Z |
| Create Time | 2025-01-13T15:17:32Z |
| Citation Date | 2024-07-15 |
| File Access Request | True |
Export metadata
Static metadata exports available for this published dataset version:
Complete Dataverse metadata
Expected crawler behaviour
Use a stable, truthful User-Agent with product/version and a working contact URL. Across all IP addresses and HTTP connections used by one crawler identity, allow no more than 5 requests in flight and wait at least 20 seconds between request starts. Crawl URLs listed in the catalog sitemap, including file pages and download URLs when they are published, use conditional requests, honor Retry-After, and apply exponential backoff after errors.
The welcome page may link to the interactive repository for human navigation. Automated clients must not treat that human link as a catalog crawl target.
Read the live machine-readable crawler policy before and during a crawl. Stop crawling when it reports CPU or memory utilization at or above 80% and 80% respectively.
