This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| from __future__ import absolute_import | |
| from __future__ import division | |
| from __future__ import print_function | |
| import gzip | |
| import os | |
| import sys | |
| import time | |
| from six.moves import urllib |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| import tensorflow as tf | |
| import numpy as np | |
| import random | |
| import time | |
| from tensorflow.python.framework import ops | |
| from tensorflow.python.framework import dtypes | |
| dataset_path = "/cs/home/is59/database/" | |
| train_labels_file = "labels.csv" |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| import tensorflow as tf | |
| import numpy as np | |
| import random | |
| import time | |
| from tensorflow.python.framework import ops | |
| from tensorflow.python.framework import dtypes | |
| dataset_path = "/cs/home/is59/database/" | |
| train_labels_file = "labels.csv" |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| """Reproduce the SwissDecisionSummaryTranslations val/test split bug. | |
| The released headnote test split (2,000 BGEs) is almost entirely BGE volumes | |
| 101-108 (1975-1982) instead of the recency-based split described in the | |
| SwiLTra-Bench paper. Cause: at the code state that generated the data | |
| (JoelNiklaus/SwissLegalTranslations @ 761fc5c8, 2024-11-21), get_split_cols | |
| sorted twice: | |
| df = df.sort_values(by='year', ascending=False) # intended recency | |
| ... |