Evaluating Dialogue Systems via an Opinion
Open this dataset in the live repository
- Persistent identifier
- doi:10.60507/FK2/FX37GD
- Published version
- 1.0
- Publication date
- 2023-06-05
- License
- CC0 1.0
Description
Dialogue systems are a significant field of research and development in artificial intelligence. Until today, the evaluation of such algorithms happens in one fundamental way. They solve "hypothetical problems," i.e., dialog systems are tested by being asked to respond in specific scenarios and provide a "solution", i.e. a reply, to a "problem". The replies are either explicitly compared to reference responses using overlap-based metrics (e.g., BLEU) or are evaluated by human annotators, which can also be seen as compared to (implicit) references. Instead, we propose to ask dialogue systems to tell whether a sample "solution" to a sample "problem" is good or bad. In other words, we ask another dialog system whether a conversation is fluent and coherent or not, and to what degree. In our experiments, we show how to evaluate dialogue systems by "asking for an opinion" and that it indeed offers an additional perspective on assessing these methods.
Creators
- Nedelchev, Rostislav
Keywords
dialogue systems, evaluation, opinion
Files
| File | Type | Bytes | Checksum |
|---|---|---|---|
| hred.py | text/x-python | 7938 | MD5 0e2689540b36a37c37c976aa893cae7b |
| embedding_metric.py | text/x-python | 1774 | MD5 06f2e4c930bec0c35599a3b14c005f86 |
| convai1_preprocess.py | text/x-python | 6445 | MD5 34badcfeeb6e613da58e01069c69d0ba |
| samples.txt | text/plain | 1094130 | MD5 4418097cfd14bc24c14768a2009cf5c8 |
| convai2.sqlite | application/octet-stream | 1130496 | MD5 99054ec8892a03307266f5f346f43470 |
| loss.py | text/x-python | 1680 | MD5 93bbfc81f386b03e5f767de039b8a608 |
| sqlitedict_compress.py | text/x-python | 223 | MD5 973f8bc6bbdd46f8d2745a9eb47c0a00 |
| conversation_length.pkl | application/octet-stream | 16644 | MD5 173d6c89c88d3ded02f72392703bb9fd |
| sentences.pkl | application/octet-stream | 2943221 | MD5 445adabc6deb87959aea175b54564346 |
| config.txt | text/plain | 2086 | MD5 3d22574ac31c49af8c47282a364aba3e |
| gpt_convai_eval.py | text/x-python | 7666 | MD5 84e5e442eb3ccbf11b2f26c43794404a |
| sentences.pkl | application/octet-stream | 5279906 | MD5 8dbc7e678bd804380d3acfa51e5b064f |
| probability.py | text/x-python | 860 | MD5 2ac32870e25170b83038aefd9190a5a9 |
| encoder.py | text/x-python | 8097 | MD5 882f7fa4798f95d1eeb9a1db788f9d5b |
| convai1.json | application/json | 5804611 | MD5 630c9b1f776b09e0be71290747bc635e |
| convai1.sqlite | application/octet-stream | 856064 | MD5 1a3ecfaa638e087a7af3455eaad2e3aa |
| events.out.tfevents.1604766709.ff8fc1a53fe2 | application/octet-stream | 7540 | MD5 935c39fec78d21f628f66ae076e0288f |
| convai2.sqlite | application/octet-stream | 1134592 | MD5 3921916065420989e5b87254f53b3adb |
| samples.txt | text/plain | 1017054 | MD5 7c4c004fe1cd7d8ae81aaf093fff42f9 |
| eval_embed.py | text/x-python | 1138 | MD5 8d9fd5ab381a75d1ee374ff93d93d33c |
| conversation_length.pkl | application/octet-stream | 4954 | MD5 a37da43749e1b11cfd4026e5531c6634 |
| vhred.py | text/x-python | 12544 | MD5 0c63debba194f0dbaccd7dae9a259342 |
| ubuntu_preprocess.py | text/x-python | 8853 | MD5 08b4e5640f8720240393c50be9d313d5 |
| configs.py | text/x-python | 6870 | MD5 02a1ad2a0033f1195b96f2aed2cce16d |
| sentences.pkl | application/octet-stream | 5250859 | MD5 8eef179efcbc3135c72eb73e8c7b34d4 |
| word2id.pkl | application/octet-stream | 299497 | MD5 5d7cba1b2d8c6d8d5ba764d6dad4c533 |
| vhcr.py | text/x-python | 19884 | MD5 b8ce2c26b2334803a6ee765e2d45a7a8 |
| __init__.py | text/x-python | 147 | MD5 231a8ea4ade1ee4b2aa93a84ad62897c |
| movie_titles_metadata.txt | text/plain | 67289 | MD5 e7e6beb7c5b5486d21a6bdee314618d5 |
| sentences.pkl | application/octet-stream | 2542718 | MD5 a8a963b0e96c9e4c7beb0b400a57e85e |
| config.txt | text/plain | 2077 | MD5 68681f9732c66ed3300ff553292f2ee2 |
| decoder.py | text/x-python | 12223 | MD5 99d0bc27da6b3b5aac146de1c5c5d486 |
| 30.pkl | application/octet-stream | 128405303 | MD5 3c121e1beb0332b1de983fccf3f2f9b7 |
| solver.py | text/x-python | 38317 | MD5 c9c217e4152b9e4d47d33d75434e5dcc |
| samples.txt | text/plain | 986246 | MD5 c5706925d8621fee7bd76964c54d43eb |
| mask.py | text/x-python | 1207 | MD5 3b6813acd683ac6b93aab9574eba257c |
| rnncells.py | text/x-python | 2526 | MD5 a602b7f2644eb12a62cc3383f9d22de6 |
| ae_mapper.py | text/x-python | 6015 | MD5 bf35be8d01a2579a16fae7c0c4759c78 |
| 30.pkl | application/octet-stream | 296152508 | MD5 e7e1bfc10b05521fc2559e8d442ef642 |
| chameleons.pdf | application/pdf | 290691 | MD5 7a24189a5b18066874a458b5f928ef07 |
| config.txt | text/plain | 2080 | MD5 361459e347e310f59e898dddb2a4aa6d |
| __init__.py | text/x-python | 259 | MD5 4166b9cc6758907c60dc4c7316938a90 |
| sentences.pkl | application/octet-stream | 42099789 | MD5 80ce0039071b670266512c81191ee909 |
| time_track.py | text/x-python | 1132 | MD5 c6948eda97da898d8424ccdfb2cab2a6 |
| 30.pkl | application/octet-stream | 186459282 | MD5 10ab38bc59bc3376d3f608abd61e1481 |
| movie_lines.txt | text/plain | 34641919 | MD5 5090b93272d88d91681710cb78ae8e2d |
| __init__.py | text/x-python | 0 | MD5 d41d8cd98f00b204e9800998ecf8427e |
| cornell_preprocess.py | text/x-python | 8548 | MD5 f66b71fce64cd13e188a039df200c119 |
| __init__.py | text/x-python | 273 | MD5 c7ef12eb8ebe665435fc81a370615968 |
| eval_convai.py | text/x-python | 1688 | MD5 3a20fc9b13dca4a13bd47d7df15958ac |
| vocab.py | text/x-python | 4795 | MD5 1a43e9cc67888a12aa809e45ca26238f |
| convai2.sqlite | application/octet-stream | 1146880 | MD5 22f2249347a80914b2d9fd3239716764 |
| convai1.sqlite | application/octet-stream | 856064 | MD5 3bc7080bab23c9bc3229bea6f89fa685 |
| events.out.tfevents.1604773321.ff8fc1a53fe2 | application/octet-stream | 7540 | MD5 cf0763b3851aca0a118c3b18db23c52b |
| README.txt | text/plain | 4182 | MD5 9671d129f173eb13a2064f6e2a996518 |
| data_loader.py | text/x-python | 2743 | MD5 be813c6dfea1a7c71c5bce42305a828d |
| conversation_length.pkl | application/octet-stream | 4320 | MD5 e0183eaa335390c3a1758f82f6fd69b6 |
| events.out.tfevents.1604756310.ff8fc1a53fe2 | application/octet-stream | 3070 | MD5 e7ada1fbc7994ceae47ffaf84e627df5 |
| 30.pkl | application/octet-stream | 212482594 | MD5 5248334cb6f456a8b35adf1a1599453e |
| sentence_length.pkl | application/octet-stream | 124157 | MD5 76437aef3170319c05a4ab818e8b8996 |
| train.py | text/x-python | 1665 | MD5 43192d2528f26c287e0aee575386b181 |
| events.out.tfevents.1604760827.ff8fc1a53fe2 | application/octet-stream | 3070 | MD5 40975185bea819107af55b375df9f766 |
| convai1.sqlite | application/octet-stream | 860160 | MD5 006c8e345788b40732a51c390826f7b9 |
| eval_4models_convai.sh | application/x-sh | 1179 | MD5 7aa14c8e2311ce7179887e4acff5d3dd |
| movie_conversations.txt | text/plain | 6760930 | MD5 0d5937e4989e9d797b6f7b2fc14ae208 |
| train_4models.sh | application/x-sh | 210 | MD5 2a7673061c4784d39c3dd421dbd58cbf |
| sentence_length.pkl | application/octet-stream | 123603 | MD5 c007e52a67f6d7e8d928117c97e3aba3 |
| eval.py | text/x-python | 1168 | MD5 8e5f271adb9a42b238b5be36a14021f9 |
| id2word.pkl | application/octet-stream | 299497 | MD5 6963378392fc52c6a4681f4771a55737 |
| conversation_length.pkl | application/octet-stream | 16644 | MD5 7aaa3d201d1d05c6f63762a3c5492717 |
| sentence_length.pkl | application/octet-stream | 60871 | MD5 f584263389cafecab4dcba99ad73b1e7 |
| env.yml | application/octet-stream | 7745 | MD5 5f3c5c7012ebfa089a06d7850b3b8e01 |
| convert.py | text/x-python | 1922 | MD5 dba8122c3a05f62dfe410b91ca3fce7d |
| convai1_results.pickle.bz2 | application/x-bzip2 | 13323619 | MD5 e5dffeba0b7e5284121d496d369a81a0 |
| seq2seq.py | text/x-python | 3668 | MD5 d12b1b55f5a3f21789e1e5b4b5640cb6 |
| convai2_preprocess.py | text/x-python | 6460 | MD5 c940529e00003c4a9667620176b293b2 |
| sentence_length.pkl | application/octet-stream | 52383 | MD5 b67d7f200264cdf6a44ca52cae842e31 |
| conversation_length.pkl | application/octet-stream | 133094 | MD5 c96b779c685ef60f9cf2b8d0808b3528 |
| convai2.json | application/json | 6636788 | MD5 b3161de7cdc5f7f773970983c89082d6 |
| correlations.ipynb | application/x-ipynb+json | 71458 | MD5 f9453dd3a93a4f1c8f293396db149516 |
| pad.py | text/x-python | 953 | MD5 d8c1f4e5cda0410d0a9faa40a1937d8a |
| convai2_results.pickle.bz2 | application/x-bzip2 | 21374451 | MD5 ace2a7d6e305fa91644e76faa945265f |
| samples.txt | text/plain | 1095517 | MD5 d53c15ca25c8aab41e90842dfcbd91ab |
| convai1.sqlite | application/octet-stream | 868352 | MD5 dc77dc296d56b2ecb3396a302935e4c3 |
| beam_search.py | text/x-python | 6036 | MD5 55786059366fd0e2415ac61230070465 |
| tensorboard.py | text/x-python | 895 | MD5 b7e500f2e8d5c0a108a33b0c02881b0d |
| convai2.sqlite | application/octet-stream | 1134592 | MD5 6f07962b1939d4913748bafd0a5f8c85 |
| tokenizer.py | text/x-python | 2127 | MD5 41e72881e530131faf0c44ffedf230f4 |
| feedforward.py | text/x-python | 886 | MD5 3524809e47bc22e83f7cc5ef398e86ae |
| raw_script_urls.txt | text/plain | 56177 | MD5 7ebdcd4f5963f33ad2b9376ee3f37f01 |
| movie_characters_metadata.txt | text/plain | 705695 | MD5 d99068087e5f217e198f42d973ff4b07 |
| config.txt | text/plain | 2077 | MD5 bbde8eaf244b11b38e4af9bc4989fc67 |
| sentence_length.pkl | application/octet-stream | 997403 | MD5 84356fa01d52ceb9e8aee023e8e307d6 |
| bow.py | text/x-python | 1102 | MD5 aa7a0d06cae20c30180aeb0d426e184c |
| Readme.md | text/markdown | 2000 | MD5 91148f3145562aa87cb6db5021ed2003 |
Citation
Nedelchev, Rostislav, 2023-06-05, Evaluating Dialogue Systems via an Opinion, doi:10.60507/FK2/FX37GD, V1.0
Additional Dataverse fields
| Id | 3 |
|---|---|
| Dataset Type | dataset |
| Internal Version Number | 13 |
| Latest Version Publishing State | RELEASED |
| Deaccession Link | Not supplied |
| Release Time | 2023-06-05T07:58:51Z |
| Create Time | 2023-05-20T10:41:41Z |
| Citation Date | 2023-06-05 |
| File Access Request | True |
Export metadata
Static metadata exports available for this published dataset version:
Complete Dataverse metadata
Expected crawler behaviour
Use a stable, truthful User-Agent with product/version and a working contact URL. Across all IP addresses and HTTP connections used by one crawler identity, allow no more than 5 requests in flight and wait at least 20 seconds between request starts. Crawl URLs listed in the catalog sitemap, including file pages and download URLs when they are published, use conditional requests, honor Retry-After, and apply exponential backoff after errors.
The welcome page may link to the interactive repository for human navigation. Automated clients must not treat that human link as a catalog crawl target.
Read the live machine-readable crawler policy before and during a crawl. Stop crawling when it reports CPU or memory utilization at or above 80% and 80% respectively.
