The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
Error code: DatasetGenerationCastError
Exception: DatasetGenerationCastError
Message: An error occurred while generating the dataset
All the data files must have the same columns, but at some point there are 1276 new columns ({'VOPP1', 'CST7', 'CCDC39', 'GIMAP7', 'SERHL2', 'SPOCK2', 'LAMB3', 'LGI4', 'PTGS1', 'SKP1', 'MSLN', 'IL6', 'EEF1D', 'NOTCH1', 'FCGR1A', 'CD79B', 'RBFOX3', 'SOX18', 'INSL5', 'CCL2', 'FADD', 'CDKN1A', 'ARPC5', 'SMS', 'CALU', 'PLCE1', 'STC2', 'IGHGP', 'CTSB', 'CYP4B1', 'VSIG4', 'CDKN2A', 'EBPL', 'CMA1', 'KDR', 'ABCC8', 'IL4', 'S100P', 'SLC4A1', 'CES1', 'SCGB2A1', 'DUSP2', 'TREM2', 'MMP3', 'PCDH10', 'FSCN1', 'NOTCH2', 'GSTA1', 'WNT2', 'ENAH', 'FASN', 'SFXN1', 'TYROBP', 'VAMP8', 'AREG', 'CDK6', 'ADIPOQ', 'BIRC3', 'REXO4', 'C15orf48', 'IFNGR1', 'SLC25A37', 'SMAD3', 'APOC1', 'CTHRC1', 'FHL1', 'SPI1', 'SCG2', 'GKN2', 'OTOP2', 'CACNG4', 'MS4A6A', 'NDUFA4L2', 'ICOSLG', 'IL4I1', 'TENT5C', 'RUNX1T1', 'TNFRSF25', 'PTGDS', 'ITM2C', 'LAMA2', 'NPM3', 'VEGFA', 'DPP6', 'TMIGD1', 'ENTPD1', 'MTRNR2L11', 'IL22RA2', 'ODF2L', 'CD40', 'RSPO3', 'CTSE', 'MARCO', 'FOS', 'COL8A1', 'ECSCR', 'ADAM17', 'APOA5', 'DCLK1', 'RPS4Y1', 'MYO5B', 'NKG7', 'ASCL1', 'SST', 'CEACAM1', 'INSM1', 'TP73', 'DEPP1', 'CYP3A4', 'TXLNA', 'FGL2', 'COTL1', 'SERPINA1', 'TFPI', 'ICA1', 'HAMP', 'MMRN2', 'CYP2A7', 'CCR3', 'GZMB', 'PIM2', 'NF1', 'IFNL1', 'RHOA', 'COL19A1', 'PMP22', 'CXCL5', 'FOXP3', 'RNF43', 'KLK11', 'DSP', 'ASAH1', 'CCL26', 'FOXC2', 'LGR6', 'EDN1', 'OGN', 'ANKRD28', 'CXCR5', 'GPR34', 'CD34', 'CCPG1', 'ID2', 'HPX', 'CD276', 'AQP2', 'F3', 'CDCA7', 'MS4A7', 'COL5A2', 'CYP2F1', 'DAPK3', 'NXPH1', 'CD2', 'TKT', 'SOX17', 'TRAC', 'GPR183', 'SHANK3', 'PCOLCE2', 'MCEMP1', 'SVIL', 'TK1', 'SLC22A8', 'CA2', 'CDKN2C', 'THAP2', 'S
...
'FCN3', 'CDHR5', 'FASLG', 'HSP90B1', 'ORC6', 'NPC2', 'HDC', 'LRRC15', 'IGF1R', 'CTTN', 'SEC61B', 'EPOR', 'CDH1', 'LILRB4', 'NOTCH3', 'ANXA3', 'APCDD1', 'ID4', 'TRDN', 'SLC7A11', 'EIF4EBP1', 'MYBPC1', 'MYOM1', 'IFIT3', 'DUSP5', 'SLC26A2', 'GDF15', 'CDX1', 'FFAR4', 'FOXI1', 'GUCA2B', 'SEC24A', 'SLC29A4', 'GZMK', 'ITGAM', 'PCNA', 'LILRA5', 'EDNRB', 'CDK15', 'EHF', 'PTPRB', 'BAALC', 'SCG5', 'CLEC10A', 'POSTN', 'FGFBP1', 'UPK1B', 'ELOVL5', 'PDK4', 'EPHA2', 'HOXD9', 'GRB14', 'TRAT1', 'PGR', 'SDC1', 'CX3CR1', 'ABCC11', 'GATM', 'TP53', 'HRCT1', 'BEST4', 'XCR1', 'ATM', 'TAGLN', 'CEBPB', 'KRT20', 'SOD3', 'NXPE4', 'CLEC14A', 'FZD7', 'LAG3', 'GNLY', 'C1QB', 'FCN2', 'CP', 'NID1', 'LY6E', 'EMP3', 'STAT2', 'ACTG2', 'CD247', 'ICAM1', 'CSF2RA', 'RGS5', 'PPP1R1B', 'VEGFC', 'CCL21', 'TMEM61', 'COCH', 'S100A4', 'STAT4', 'ARPC3', 'CDC42EP1', 'PIM1', 'IL12RB2', 'MEIS2', 'SDC4', 'MFAP5', 'FEZ1', 'RTN4', 'SLC12A2', 'S100B', 'PRF1', 'SSR2', 'JCHAIN', 'BAMBI', 'ACE2', 'NOSTRIN', 'PELI1', 'PLK1', 'ACACB', 'TFAP2A', 'CCL15', 'FCGR2A', 'GREM1', 'C20orf85', 'S100A1', 'GPC3', 'CSTA', 'FCER1A', 'OSTC', 'MYH14', 'C2orf88', 'SMIM14', 'TPSG1', 'IL12B', 'MYO6', 'C11orf96', 'NT5E', 'CILP', 'RARRES1', 'SFTPD', 'TRBC1', 'CD3D', 'WARS', 'GZMH', 'TUBA1A', 'ETV5', 'TCEAL7', 'TGFBR1', 'NOD1', 'VTN', 'CEACAM7', 'GPX3', 'TRAPPC3', 'C7', 'CTSL', 'GEM', 'PTTG1', 'MRC1', 'APOBEC3A', 'PBK', 'SEC62', 'ID1', 'IGFBP3', 'LDHB', 'MMP11', 'EPHA4', 'CRYBA2', 'PABPC1', 'IL15RA', 'NAT8', 'CFB', 'CD24', 'CLCA2', 'DNASE1L3', 'RGS16'}) and 3 missing columns ({'y_coord', 'cell_id', 'x_coord'}).
This happened while the csv dataset builder was generating data using
hf://datasets/GHISTPlus/GHIST-Plus-bundle/evaluation_data/atera/cell_gene_matrix_filtered.csv (at revision 0932796efd6a92fcc9c4bfeaf447a89ccce51269), ['hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/atera/cell_coords.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/atera/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_coords.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_type_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast5k/cell_coords_histology_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast5k/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/imputation/figure3_coordinates.csv.gz', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/imputation/figure3_ground_truth.csv.gz']
Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)
Traceback: Traceback (most recent call last):
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1848, in _prepare_split_single
writer.write_table(table)
~~~~~~~~~~~~~~~~~~^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 765, in write_table
self._write_table(pa_table, writer_batch_size=writer_batch_size)
~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 773, in _write_table
pa_table = table_cast(pa_table, self._schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2378, in table_cast
return cast_table_to_schema(table, schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2306, in cast_table_to_schema
raise CastError(
...<3 lines>...
)
datasets.table.CastError: Couldn't cast
Unnamed: 0: int64
A2M: int64
ABCA8: int64
ABCC8: int64
ABCC11: int64
ACACB: int64
ACE: int64
ACE2: int64
ACKR1: int64
ACTA2: int64
ACTB: int64
ACTG2: int64
ACTN1: int64
ADAM9: int64
ADAM17: int64
ADAM28: int64
ADAMTS1: int64
ADGRE1: int64
ADGRE5: int64
ADGRL4: int64
ADH1B: int64
ADH1C: int64
ADH4: int64
ADIPOQ: int64
ADRA2A: int64
AEBP1: int64
AFAP1L2: int64
AGER: int64
AGR3: int64
AGTR1: int64
AHSP: int64
AIF1: int64
AIRE: int64
AKR1C1: int64
AKR1C3: int64
AKR7A3: int64
AKT1: int64
ALAS2: int64
ALDH1A3: int64
ALDH1B1: int64
ALOX5AP: int64
AMY2A: int64
ANGPT2: int64
ANK2: int64
ANKRD28: int64
ANKRD29: int64
ANKRD30A: int64
ANO7: int64
ANPEP: int64
ANXA1: int64
ANXA3: int64
ANXA13: int64
APC: int64
APCDD1: int64
APOA5: int64
APOB: int64
APOBEC3A: int64
APOBEC3B: int64
APOC1: int64
APOD: int64
APOE: int64
APOLD1: int64
AQP1: int64
AQP2: int64
AQP3: int64
AQP8: int64
AQP9: int64
AR: int64
AREG: int64
ARFGEF3: int64
ARG1: int64
ARHGAP24: int64
ARID1A: int64
ARL14: int64
ARPC3: int64
ARPC5: int64
ARSG: int64
ARX: int64
ASAH1: int64
ASCL1: int64
ASCL2: int64
ASCL3: int64
ASPN: int64
ATM: int64
ATOH1: int64
ATP1B1: int64
ATP5F1B: int64
ATP5MC2: int64
ATP5MD: int64
AVIL: int64
AVPR1A: int64
AZGP1: int64
B3GNT6: int64
BAALC: int64
BACE2: int64
BAIAP2L1: int64
BAMBI: int64
BANK1: int64
BASP1: int64
BATF: int64
BATF3: int64
BBOX1: int64
BCAS1: int64
BCL2: int64
BCL2L11: int64
BEST2: int64
BEST4: int64
BIRC3: int64
BMP4: int64
BMP5: int64
BMX: int64
BRAF: int64
BRCA2: int64
BTF3: int64
B
...
int64
THBS1: int64
THBS2: int64
THY1: int64
TIFA: int64
TIGIT: int64
TIMP3: int64
TIMP4: int64
TK1: int64
TKT: int64
TM4SF4: int64
TM4SF18: int64
TMA7: int64
TMBIM6: int64
TMC5: int64
TMEM52B: int64
TMEM61: int64
TMEM100: int64
TMEM147: int64
TMEM174: int64
TMIGD1: int64
TMPRSS2: int64
TNC: int64
TNF: int64
TNFAIP3: int64
TNFRSF1B: int64
TNFRSF9: int64
TNFRSF13B: int64
TNFRSF13C: int64
TNFRSF17: int64
TNFRSF18: int64
TNFRSF25: int64
TNFSF13B: int64
TNS4: int64
TNXB: int64
TOMM7: int64
TOP2A: int64
TOX: int64
TP53: int64
TP63: int64
TP73: int64
TPD52: int64
TPSAB1: int64
TPSG1: int64
TRAC: int64
TRAF4: int64
TRAPPC3: int64
TRAT1: int64
TRBC1: int64
TRBC2: int64
TRDN: int64
TRDV1: int64
TREM2: int64
TRGV4: int64
TRH: int64
TRIB1: int64
TRPC6: int64
TRPM5: int64
TSPAN8: int64
TSPAN19: int64
TTR: int64
TUBA1A: int64
TUBA1B: int64
TUBA4A: int64
TUBB: int64
TUBB2B: int64
TXLNA: int64
TYMS: int64
TYROBP: int64
UBD: int64
UBE2C: int64
UCN3: int64
UCP1: int64
UGP2: int64
UGT2A3: int64
UGT2B17: int64
UMOD: int64
UPK1B: int64
UPK3B: int64
UQCC2: int64
USP53: int64
VAMP8: int64
VCAN: int64
VEGFA: int64
VEGFC: int64
VOPP1: int64
VPREB3: int64
VSIG4: int64
VSIR: int64
VTN: int64
VWA5A: int64
VWA5B2: int64
VWF: int64
WARS: int64
WFDC2: int64
WFS1: int64
WNT2: int64
WNT5B: int64
WT1: int64
XBP1: int64
XCL2: int64
XCR1: int64
YAF2: int64
ZEB1: int64
ZEB2: int64
ZNF562: int64
ZNF683: int64
-- schema metadata --
pandas: '{"index_columns": [{"kind": "range", "name": null, "start": 0, "' + 136663
to
{'cell_id': Value('int64'), 'x_coord': Value('float64'), 'y_coord': Value('float64')}
because column names don't match
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
~~~~~~~~~~~~~~~~~~~~~~~~~^
builder, max_dataset_size_bytes=max_dataset_size_bytes
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1694, in _prepare_split
for job_id, done, content in self._prepare_split_single(
~~~~~~~~~~~~~~~~~~~~~~~~~~^
gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
):
^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1850, in _prepare_split_single
raise DatasetGenerationCastError.from_cast_error(
...<4 lines>...
)
datasets.exceptions.DatasetGenerationCastError: An error occurred while generating the dataset
All the data files must have the same columns, but at some point there are 1276 new columns ({'VOPP1', 'CST7', 'CCDC39', 'GIMAP7', 'SERHL2', 'SPOCK2', 'LAMB3', 'LGI4', 'PTGS1', 'SKP1', 'MSLN', 'IL6', 'EEF1D', 'NOTCH1', 'FCGR1A', 'CD79B', 'RBFOX3', 'SOX18', 'INSL5', 'CCL2', 'FADD', 'CDKN1A', 'ARPC5', 'SMS', 'CALU', 'PLCE1', 'STC2', 'IGHGP', 'CTSB', 'CYP4B1', 'VSIG4', 'CDKN2A', 'EBPL', 'CMA1', 'KDR', 'ABCC8', 'IL4', 'S100P', 'SLC4A1', 'CES1', 'SCGB2A1', 'DUSP2', 'TREM2', 'MMP3', 'PCDH10', 'FSCN1', 'NOTCH2', 'GSTA1', 'WNT2', 'ENAH', 'FASN', 'SFXN1', 'TYROBP', 'VAMP8', 'AREG', 'CDK6', 'ADIPOQ', 'BIRC3', 'REXO4', 'C15orf48', 'IFNGR1', 'SLC25A37', 'SMAD3', 'APOC1', 'CTHRC1', 'FHL1', 'SPI1', 'SCG2', 'GKN2', 'OTOP2', 'CACNG4', 'MS4A6A', 'NDUFA4L2', 'ICOSLG', 'IL4I1', 'TENT5C', 'RUNX1T1', 'TNFRSF25', 'PTGDS', 'ITM2C', 'LAMA2', 'NPM3', 'VEGFA', 'DPP6', 'TMIGD1', 'ENTPD1', 'MTRNR2L11', 'IL22RA2', 'ODF2L', 'CD40', 'RSPO3', 'CTSE', 'MARCO', 'FOS', 'COL8A1', 'ECSCR', 'ADAM17', 'APOA5', 'DCLK1', 'RPS4Y1', 'MYO5B', 'NKG7', 'ASCL1', 'SST', 'CEACAM1', 'INSM1', 'TP73', 'DEPP1', 'CYP3A4', 'TXLNA', 'FGL2', 'COTL1', 'SERPINA1', 'TFPI', 'ICA1', 'HAMP', 'MMRN2', 'CYP2A7', 'CCR3', 'GZMB', 'PIM2', 'NF1', 'IFNL1', 'RHOA', 'COL19A1', 'PMP22', 'CXCL5', 'FOXP3', 'RNF43', 'KLK11', 'DSP', 'ASAH1', 'CCL26', 'FOXC2', 'LGR6', 'EDN1', 'OGN', 'ANKRD28', 'CXCR5', 'GPR34', 'CD34', 'CCPG1', 'ID2', 'HPX', 'CD276', 'AQP2', 'F3', 'CDCA7', 'MS4A7', 'COL5A2', 'CYP2F1', 'DAPK3', 'NXPH1', 'CD2', 'TKT', 'SOX17', 'TRAC', 'GPR183', 'SHANK3', 'PCOLCE2', 'MCEMP1', 'SVIL', 'TK1', 'SLC22A8', 'CA2', 'CDKN2C', 'THAP2', 'S
...
'FCN3', 'CDHR5', 'FASLG', 'HSP90B1', 'ORC6', 'NPC2', 'HDC', 'LRRC15', 'IGF1R', 'CTTN', 'SEC61B', 'EPOR', 'CDH1', 'LILRB4', 'NOTCH3', 'ANXA3', 'APCDD1', 'ID4', 'TRDN', 'SLC7A11', 'EIF4EBP1', 'MYBPC1', 'MYOM1', 'IFIT3', 'DUSP5', 'SLC26A2', 'GDF15', 'CDX1', 'FFAR4', 'FOXI1', 'GUCA2B', 'SEC24A', 'SLC29A4', 'GZMK', 'ITGAM', 'PCNA', 'LILRA5', 'EDNRB', 'CDK15', 'EHF', 'PTPRB', 'BAALC', 'SCG5', 'CLEC10A', 'POSTN', 'FGFBP1', 'UPK1B', 'ELOVL5', 'PDK4', 'EPHA2', 'HOXD9', 'GRB14', 'TRAT1', 'PGR', 'SDC1', 'CX3CR1', 'ABCC11', 'GATM', 'TP53', 'HRCT1', 'BEST4', 'XCR1', 'ATM', 'TAGLN', 'CEBPB', 'KRT20', 'SOD3', 'NXPE4', 'CLEC14A', 'FZD7', 'LAG3', 'GNLY', 'C1QB', 'FCN2', 'CP', 'NID1', 'LY6E', 'EMP3', 'STAT2', 'ACTG2', 'CD247', 'ICAM1', 'CSF2RA', 'RGS5', 'PPP1R1B', 'VEGFC', 'CCL21', 'TMEM61', 'COCH', 'S100A4', 'STAT4', 'ARPC3', 'CDC42EP1', 'PIM1', 'IL12RB2', 'MEIS2', 'SDC4', 'MFAP5', 'FEZ1', 'RTN4', 'SLC12A2', 'S100B', 'PRF1', 'SSR2', 'JCHAIN', 'BAMBI', 'ACE2', 'NOSTRIN', 'PELI1', 'PLK1', 'ACACB', 'TFAP2A', 'CCL15', 'FCGR2A', 'GREM1', 'C20orf85', 'S100A1', 'GPC3', 'CSTA', 'FCER1A', 'OSTC', 'MYH14', 'C2orf88', 'SMIM14', 'TPSG1', 'IL12B', 'MYO6', 'C11orf96', 'NT5E', 'CILP', 'RARRES1', 'SFTPD', 'TRBC1', 'CD3D', 'WARS', 'GZMH', 'TUBA1A', 'ETV5', 'TCEAL7', 'TGFBR1', 'NOD1', 'VTN', 'CEACAM7', 'GPX3', 'TRAPPC3', 'C7', 'CTSL', 'GEM', 'PTTG1', 'MRC1', 'APOBEC3A', 'PBK', 'SEC62', 'ID1', 'IGFBP3', 'LDHB', 'MMP11', 'EPHA4', 'CRYBA2', 'PABPC1', 'IL15RA', 'NAT8', 'CFB', 'CD24', 'CLCA2', 'DNASE1L3', 'RGS16'}) and 3 missing columns ({'y_coord', 'cell_id', 'x_coord'}).
This happened while the csv dataset builder was generating data using
hf://datasets/GHISTPlus/GHIST-Plus-bundle/evaluation_data/atera/cell_gene_matrix_filtered.csv (at revision 0932796efd6a92fcc9c4bfeaf447a89ccce51269), ['hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/atera/cell_coords.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/atera/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_coords.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_type_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast5k/cell_coords_histology_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast5k/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/imputation/figure3_coordinates.csv.gz', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/imputation/figure3_ground_truth.csv.gz']
Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
cell_id int64 | x_coord float64 | y_coord float64 |
|---|---|---|
1 | 77,496.841733 | 4,082.609656 |
2 | 77,537.900401 | 4,370.606505 |
3 | 77,184.576033 | 3,873.576296 |
4 | 77,250.12938 | 3,869.166524 |
5 | 77,384.880698 | 4,028.793313 |
6 | 77,402.927754 | 4,039.828016 |
7 | 77,608.767233 | 4,698.044269 |
8 | 77,262.205517 | 4,773.734259 |
10 | 77,645.965348 | 4,964.74568 |
11 | 77,504.164625 | 4,436.79834 |
12 | 77,533.368381 | 4,458.082255 |
13 | 77,377.910075 | 4,424.744235 |
14 | 77,405.112541 | 4,455.038862 |
15 | 77,491.16299 | 4,457.368883 |
16 | 77,757.498847 | 4,406.618391 |
17 | 77,681.642409 | 4,687.665622 |
18 | 77,767.367421 | 4,676.995191 |
19 | 77,661.391378 | 4,519.183994 |
20 | 77,675.221637 | 4,090.759088 |
21 | 77,707.766352 | 4,350.739595 |
22 | 77,670.613108 | 4,131.705586 |
23 | 75,576.460471 | 5,207.854871 |
24 | 72,702.227606 | 4,879.310757 |
25 | 71,615.779445 | 4,955.108398 |
27 | 71,881.301518 | 4,835.285078 |
28 | 73,490.032672 | 4,964.085858 |
29 | 72,101.87823 | 4,718.409513 |
30 | 74,586.200639 | 4,941.230864 |
31 | 73,822.263766 | 4,986.84304 |
32 | 73,926.958675 | 5,022.137536 |
33 | 74,305.788181 | 4,863.563015 |
34 | 74,314.140861 | 4,908.361581 |
36 | 72,131.051938 | 5,246.706835 |
37 | 74,577.292972 | 5,265.406334 |
38 | 72,293.407026 | 5,231.636427 |
39 | 73,952.353312 | 5,220.429419 |
40 | 73,816.456741 | 5,189.144562 |
41 | 74,731.157576 | 5,226.575805 |
42 | 74,768.826378 | 5,199.801215 |
43 | 74,902.167273 | 5,206.50993 |
44 | 75,432.08101 | 4,879.138671 |
45 | 75,269.804428 | 4,755.513862 |
47 | 76,934.155068 | 4,929.299145 |
48 | 76,532.563511 | 4,782.766704 |
49 | 76,625.775811 | 4,705.831506 |
50 | 76,925.040744 | 5,072.876723 |
52 | 75,675.618704 | 4,985.018186 |
53 | 75,700.048299 | 4,955.589336 |
54 | 75,679.760741 | 4,838.358601 |
55 | 76,360.769952 | 4,974.524822 |
56 | 76,035.965887 | 5,090.526428 |
57 | 76,264.704636 | 5,153.247145 |
58 | 75,707.956881 | 4,638.28954 |
60 | 75,787.707225 | 4,672.027241 |
63 | 76,622.848298 | 4,638.162573 |
64 | 77,227.710211 | 5,124.79241 |
65 | 77,017.890736 | 5,189.001011 |
66 | 77,140.003196 | 5,199.766397 |
67 | 76,903.310249 | 5,122.009091 |
68 | 77,547.153788 | 5,146.891362 |
69 | 77,254.809826 | 5,044.80904 |
70 | 77,094.931419 | 5,000.251794 |
71 | 77,099.166428 | 4,726.981833 |
72 | 77,181.046944 | 4,752.219104 |
73 | 77,288.648005 | 4,422.325687 |
74 | 77,183.830322 | 4,548.79255 |
75 | 77,316.311911 | 4,502.802414 |
76 | 75,746.846489 | 4,437.075111 |
77 | 74,481.789582 | 4,603.64529 |
78 | 74,740.094629 | 4,578.377749 |
79 | 74,614.797527 | 4,712.509595 |
81 | 75,326.258132 | 4,726.243822 |
82 | 75,203.259187 | 4,636.961879 |
84 | 75,432.077192 | 4,572.807333 |
85 | 75,303.960476 | 4,448.932645 |
88 | 73,385.801245 | 4,437.450125 |
89 | 72,769.643886 | 4,782.034326 |
90 | 73,117.966729 | 4,829.269108 |
91 | 73,161.121075 | 4,594.905998 |
92 | 73,166.042679 | 4,365.272054 |
93 | 75,017.656489 | 4,705.778952 |
94 | 74,830.960171 | 4,440.135673 |
95 | 74,883.415466 | 4,745.827906 |
97 | 76,810.227061 | 4,366.586414 |
98 | 76,753.596723 | 3,984.775452 |
100 | 75,644.886195 | 4,117.98209 |
101 | 74,480.035718 | 4,077.647946 |
106 | 75,086.049028 | 4,287.121509 |
107 | 75,865.171736 | 4,371.789775 |
110 | 76,368.870049 | 4,087.135858 |
113 | 72,251.736973 | 6,483.371723 |
114 | 72,210.426341 | 6,494.434017 |
115 | 72,225.502479 | 6,504.196605 |
116 | 72,223.740039 | 6,471.897379 |
117 | 72,175.801668 | 6,584.235427 |
118 | 71,974.014053 | 6,610.985756 |
119 | 71,974.541898 | 6,591.541713 |
120 | 72,174.0108 | 6,459.324372 |
121 | 71,999.141507 | 6,494.616251 |
122 | 72,190.200386 | 6,505.014207 |
GHIST+ data and model bundle
This bundle contains the model artifacts, predictions, evaluation inputs, comparison outputs, and plot-ready tables released with GHIST+. Source code is provided in that repository.
Download
hf download GHISTPlus/GHIST-Plus-bundle \
--repo-type dataset \
--local-dir bundle
The bundle is approximately 38 GB. Individual files can also be downloaded from this page.
Model checkpoints
The four released GHIST+ checkpoint files are:
- GHIST_plus/models/breast_multi/ghist_plus_breast_multi_checkpoint.pth
- GHIST_plus/models/breast_single/ghist_plus_breast_single_checkpoint.pth
- GHIST_plus/models/imputation/ghist_plus_gene_imputation_checkpoint.pth
- GHIST_plus/models/pancancer/ghist_plus_pancancer_checkpoint.pth
They exclude the frozen third-party UNI2-H encoder weights. Users must obtain UNI2-H directly from MahmoodLab/UNI2-h, accept its terms, and follow the Pretrained Checkpoints instructions in the GHIST+ README. That procedure reconstructs the complete checkpoints locally at the same filenames, so the existing inference commands and configs remain unchanged.
The bundle does not redistribute UNI2-H weights.
Use with the figure notebooks
The simplest layout is:
download-parent/
├── GHIST_plus/
└── bundle/
Start Jupyter from the code repository root. Figure2.ipynb through Figure5.ipynb will find ../bundle automatically:
cd GHIST_plus
jupyter lab
For another location, set:
export GHIST_BUNDLE_ROOT="/path/to/bundle"
jupyter lab
GHIST_BUNDLE_ROOT must point to the bundle directory, not its parent.
Contents
Figure 2: evaluation data, GHIST+ predictions, and comparison-model predictions under evaluation_data/, GHIST_plus/predictions/, and other_models/.
Figure 3: imputation inputs and predictions, including the bundled VQ/composition ablation predictions.
Figure 4: plot-ready tables under figure_data/figure4/.
Figure 5: paired PCC and coverage tables under figure_data/figure5/.
bundle/ ├── GHIST_plus/ │ ├── models/ │ └── predictions/ ├── evaluation_data/ ├── figure_data/ └── other_models/
The paths and lightweight input schemas were checked against the released Figure 2–5 notebooks.
Data and third-party terms
The bundle combines author-generated artifacts, processed public source data, and outputs from comparison methods. It therefore has no single blanket license; the Hugging Face license field is other. See THIRD_PARTY_NOTICES.md for component-specific sources, versions, attributions, and terms. Upstream terms remain applicable.
Tutorial data
tutorial.ipynb does not use this bundle as its DATA_ROOT. The tutorial requires a separately prepared GHIST data directory containing aligned H&E images, segmentation masks, nuclei metadata, and inputs for the selected training mode.
- Downloads last month
- 32