Dataset Preview
Duplicate
The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
The dataset generation failed because of a cast error
Error code:   DatasetGenerationCastError
Exception:    DatasetGenerationCastError
Message:      An error occurred while generating the dataset

All the data files must have the same columns, but at some point there are 1276 new columns ({'VOPP1', 'CST7', 'CCDC39', 'GIMAP7', 'SERHL2', 'SPOCK2', 'LAMB3', 'LGI4', 'PTGS1', 'SKP1', 'MSLN', 'IL6', 'EEF1D', 'NOTCH1', 'FCGR1A', 'CD79B', 'RBFOX3', 'SOX18', 'INSL5', 'CCL2', 'FADD', 'CDKN1A', 'ARPC5', 'SMS', 'CALU', 'PLCE1', 'STC2', 'IGHGP', 'CTSB', 'CYP4B1', 'VSIG4', 'CDKN2A', 'EBPL', 'CMA1', 'KDR', 'ABCC8', 'IL4', 'S100P', 'SLC4A1', 'CES1', 'SCGB2A1', 'DUSP2', 'TREM2', 'MMP3', 'PCDH10', 'FSCN1', 'NOTCH2', 'GSTA1', 'WNT2', 'ENAH', 'FASN', 'SFXN1', 'TYROBP', 'VAMP8', 'AREG', 'CDK6', 'ADIPOQ', 'BIRC3', 'REXO4', 'C15orf48', 'IFNGR1', 'SLC25A37', 'SMAD3', 'APOC1', 'CTHRC1', 'FHL1', 'SPI1', 'SCG2', 'GKN2', 'OTOP2', 'CACNG4', 'MS4A6A', 'NDUFA4L2', 'ICOSLG', 'IL4I1', 'TENT5C', 'RUNX1T1', 'TNFRSF25', 'PTGDS', 'ITM2C', 'LAMA2', 'NPM3', 'VEGFA', 'DPP6', 'TMIGD1', 'ENTPD1', 'MTRNR2L11', 'IL22RA2', 'ODF2L', 'CD40', 'RSPO3', 'CTSE', 'MARCO', 'FOS', 'COL8A1', 'ECSCR', 'ADAM17', 'APOA5', 'DCLK1', 'RPS4Y1', 'MYO5B', 'NKG7', 'ASCL1', 'SST', 'CEACAM1', 'INSM1', 'TP73', 'DEPP1', 'CYP3A4', 'TXLNA', 'FGL2', 'COTL1', 'SERPINA1', 'TFPI', 'ICA1', 'HAMP', 'MMRN2', 'CYP2A7', 'CCR3', 'GZMB', 'PIM2', 'NF1', 'IFNL1', 'RHOA', 'COL19A1', 'PMP22', 'CXCL5', 'FOXP3', 'RNF43', 'KLK11', 'DSP', 'ASAH1', 'CCL26', 'FOXC2', 'LGR6', 'EDN1', 'OGN', 'ANKRD28', 'CXCR5', 'GPR34', 'CD34', 'CCPG1', 'ID2', 'HPX', 'CD276', 'AQP2', 'F3', 'CDCA7', 'MS4A7', 'COL5A2', 'CYP2F1', 'DAPK3', 'NXPH1', 'CD2', 'TKT', 'SOX17', 'TRAC', 'GPR183', 'SHANK3', 'PCOLCE2', 'MCEMP1', 'SVIL', 'TK1', 'SLC22A8', 'CA2', 'CDKN2C', 'THAP2', 'S
...
 'FCN3', 'CDHR5', 'FASLG', 'HSP90B1', 'ORC6', 'NPC2', 'HDC', 'LRRC15', 'IGF1R', 'CTTN', 'SEC61B', 'EPOR', 'CDH1', 'LILRB4', 'NOTCH3', 'ANXA3', 'APCDD1', 'ID4', 'TRDN', 'SLC7A11', 'EIF4EBP1', 'MYBPC1', 'MYOM1', 'IFIT3', 'DUSP5', 'SLC26A2', 'GDF15', 'CDX1', 'FFAR4', 'FOXI1', 'GUCA2B', 'SEC24A', 'SLC29A4', 'GZMK', 'ITGAM', 'PCNA', 'LILRA5', 'EDNRB', 'CDK15', 'EHF', 'PTPRB', 'BAALC', 'SCG5', 'CLEC10A', 'POSTN', 'FGFBP1', 'UPK1B', 'ELOVL5', 'PDK4', 'EPHA2', 'HOXD9', 'GRB14', 'TRAT1', 'PGR', 'SDC1', 'CX3CR1', 'ABCC11', 'GATM', 'TP53', 'HRCT1', 'BEST4', 'XCR1', 'ATM', 'TAGLN', 'CEBPB', 'KRT20', 'SOD3', 'NXPE4', 'CLEC14A', 'FZD7', 'LAG3', 'GNLY', 'C1QB', 'FCN2', 'CP', 'NID1', 'LY6E', 'EMP3', 'STAT2', 'ACTG2', 'CD247', 'ICAM1', 'CSF2RA', 'RGS5', 'PPP1R1B', 'VEGFC', 'CCL21', 'TMEM61', 'COCH', 'S100A4', 'STAT4', 'ARPC3', 'CDC42EP1', 'PIM1', 'IL12RB2', 'MEIS2', 'SDC4', 'MFAP5', 'FEZ1', 'RTN4', 'SLC12A2', 'S100B', 'PRF1', 'SSR2', 'JCHAIN', 'BAMBI', 'ACE2', 'NOSTRIN', 'PELI1', 'PLK1', 'ACACB', 'TFAP2A', 'CCL15', 'FCGR2A', 'GREM1', 'C20orf85', 'S100A1', 'GPC3', 'CSTA', 'FCER1A', 'OSTC', 'MYH14', 'C2orf88', 'SMIM14', 'TPSG1', 'IL12B', 'MYO6', 'C11orf96', 'NT5E', 'CILP', 'RARRES1', 'SFTPD', 'TRBC1', 'CD3D', 'WARS', 'GZMH', 'TUBA1A', 'ETV5', 'TCEAL7', 'TGFBR1', 'NOD1', 'VTN', 'CEACAM7', 'GPX3', 'TRAPPC3', 'C7', 'CTSL', 'GEM', 'PTTG1', 'MRC1', 'APOBEC3A', 'PBK', 'SEC62', 'ID1', 'IGFBP3', 'LDHB', 'MMP11', 'EPHA4', 'CRYBA2', 'PABPC1', 'IL15RA', 'NAT8', 'CFB', 'CD24', 'CLCA2', 'DNASE1L3', 'RGS16'}) and 3 missing columns ({'y_coord', 'cell_id', 'x_coord'}).

This happened while the csv dataset builder was generating data using

hf://datasets/GHISTPlus/GHIST-Plus-bundle/evaluation_data/atera/cell_gene_matrix_filtered.csv (at revision 0932796efd6a92fcc9c4bfeaf447a89ccce51269), ['hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/atera/cell_coords.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/atera/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_coords.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_type_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast5k/cell_coords_histology_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast5k/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/imputation/figure3_coordinates.csv.gz', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/imputation/figure3_ground_truth.csv.gz']

Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1848, in _prepare_split_single
                  writer.write_table(table)
                  ~~~~~~~~~~~~~~~~~~^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 765, in write_table
                  self._write_table(pa_table, writer_batch_size=writer_batch_size)
                  ~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 773, in _write_table
                  pa_table = table_cast(pa_table, self._schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2378, in table_cast
                  return cast_table_to_schema(table, schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2306, in cast_table_to_schema
                  raise CastError(
                  ...<3 lines>...
                  )
              datasets.table.CastError: Couldn't cast
              Unnamed: 0: int64
              A2M: int64
              ABCA8: int64
              ABCC8: int64
              ABCC11: int64
              ACACB: int64
              ACE: int64
              ACE2: int64
              ACKR1: int64
              ACTA2: int64
              ACTB: int64
              ACTG2: int64
              ACTN1: int64
              ADAM9: int64
              ADAM17: int64
              ADAM28: int64
              ADAMTS1: int64
              ADGRE1: int64
              ADGRE5: int64
              ADGRL4: int64
              ADH1B: int64
              ADH1C: int64
              ADH4: int64
              ADIPOQ: int64
              ADRA2A: int64
              AEBP1: int64
              AFAP1L2: int64
              AGER: int64
              AGR3: int64
              AGTR1: int64
              AHSP: int64
              AIF1: int64
              AIRE: int64
              AKR1C1: int64
              AKR1C3: int64
              AKR7A3: int64
              AKT1: int64
              ALAS2: int64
              ALDH1A3: int64
              ALDH1B1: int64
              ALOX5AP: int64
              AMY2A: int64
              ANGPT2: int64
              ANK2: int64
              ANKRD28: int64
              ANKRD29: int64
              ANKRD30A: int64
              ANO7: int64
              ANPEP: int64
              ANXA1: int64
              ANXA3: int64
              ANXA13: int64
              APC: int64
              APCDD1: int64
              APOA5: int64
              APOB: int64
              APOBEC3A: int64
              APOBEC3B: int64
              APOC1: int64
              APOD: int64
              APOE: int64
              APOLD1: int64
              AQP1: int64
              AQP2: int64
              AQP3: int64
              AQP8: int64
              AQP9: int64
              AR: int64
              AREG: int64
              ARFGEF3: int64
              ARG1: int64
              ARHGAP24: int64
              ARID1A: int64
              ARL14: int64
              ARPC3: int64
              ARPC5: int64
              ARSG: int64
              ARX: int64
              ASAH1: int64
              ASCL1: int64
              ASCL2: int64
              ASCL3: int64
              ASPN: int64
              ATM: int64
              ATOH1: int64
              ATP1B1: int64
              ATP5F1B: int64
              ATP5MC2: int64
              ATP5MD: int64
              AVIL: int64
              AVPR1A: int64
              AZGP1: int64
              B3GNT6: int64
              BAALC: int64
              BACE2: int64
              BAIAP2L1: int64
              BAMBI: int64
              BANK1: int64
              BASP1: int64
              BATF: int64
              BATF3: int64
              BBOX1: int64
              BCAS1: int64
              BCL2: int64
              BCL2L11: int64
              BEST2: int64
              BEST4: int64
              BIRC3: int64
              BMP4: int64
              BMP5: int64
              BMX: int64
              BRAF: int64
              BRCA2: int64
              BTF3: int64
              B
              ...
              int64
              THBS1: int64
              THBS2: int64
              THY1: int64
              TIFA: int64
              TIGIT: int64
              TIMP3: int64
              TIMP4: int64
              TK1: int64
              TKT: int64
              TM4SF4: int64
              TM4SF18: int64
              TMA7: int64
              TMBIM6: int64
              TMC5: int64
              TMEM52B: int64
              TMEM61: int64
              TMEM100: int64
              TMEM147: int64
              TMEM174: int64
              TMIGD1: int64
              TMPRSS2: int64
              TNC: int64
              TNF: int64
              TNFAIP3: int64
              TNFRSF1B: int64
              TNFRSF9: int64
              TNFRSF13B: int64
              TNFRSF13C: int64
              TNFRSF17: int64
              TNFRSF18: int64
              TNFRSF25: int64
              TNFSF13B: int64
              TNS4: int64
              TNXB: int64
              TOMM7: int64
              TOP2A: int64
              TOX: int64
              TP53: int64
              TP63: int64
              TP73: int64
              TPD52: int64
              TPSAB1: int64
              TPSG1: int64
              TRAC: int64
              TRAF4: int64
              TRAPPC3: int64
              TRAT1: int64
              TRBC1: int64
              TRBC2: int64
              TRDN: int64
              TRDV1: int64
              TREM2: int64
              TRGV4: int64
              TRH: int64
              TRIB1: int64
              TRPC6: int64
              TRPM5: int64
              TSPAN8: int64
              TSPAN19: int64
              TTR: int64
              TUBA1A: int64
              TUBA1B: int64
              TUBA4A: int64
              TUBB: int64
              TUBB2B: int64
              TXLNA: int64
              TYMS: int64
              TYROBP: int64
              UBD: int64
              UBE2C: int64
              UCN3: int64
              UCP1: int64
              UGP2: int64
              UGT2A3: int64
              UGT2B17: int64
              UMOD: int64
              UPK1B: int64
              UPK3B: int64
              UQCC2: int64
              USP53: int64
              VAMP8: int64
              VCAN: int64
              VEGFA: int64
              VEGFC: int64
              VOPP1: int64
              VPREB3: int64
              VSIG4: int64
              VSIR: int64
              VTN: int64
              VWA5A: int64
              VWA5B2: int64
              VWF: int64
              WARS: int64
              WFDC2: int64
              WFS1: int64
              WNT2: int64
              WNT5B: int64
              WT1: int64
              XBP1: int64
              XCL2: int64
              XCR1: int64
              YAF2: int64
              ZEB1: int64
              ZEB2: int64
              ZNF562: int64
              ZNF683: int64
              -- schema metadata --
              pandas: '{"index_columns": [{"kind": "range", "name": null, "start": 0, "' + 136663
              to
              {'cell_id': Value('int64'), 'x_coord': Value('float64'), 'y_coord': Value('float64')}
              because column names don't match
              
              During handling of the above exception, another exception occurred:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
                  parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
                                                                        ~~~~~~~~~~~~~~~~~~~~~~~~~^
                      builder, max_dataset_size_bytes=max_dataset_size_bytes
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  )
                  ^
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
                  builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
                  ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1694, in _prepare_split
                  for job_id, done, content in self._prepare_split_single(
                                               ~~~~~~~~~~~~~~~~~~~~~~~~~~^
                      gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  ):
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1850, in _prepare_split_single
                  raise DatasetGenerationCastError.from_cast_error(
                  ...<4 lines>...
                  )
              datasets.exceptions.DatasetGenerationCastError: An error occurred while generating the dataset
              
              All the data files must have the same columns, but at some point there are 1276 new columns ({'VOPP1', 'CST7', 'CCDC39', 'GIMAP7', 'SERHL2', 'SPOCK2', 'LAMB3', 'LGI4', 'PTGS1', 'SKP1', 'MSLN', 'IL6', 'EEF1D', 'NOTCH1', 'FCGR1A', 'CD79B', 'RBFOX3', 'SOX18', 'INSL5', 'CCL2', 'FADD', 'CDKN1A', 'ARPC5', 'SMS', 'CALU', 'PLCE1', 'STC2', 'IGHGP', 'CTSB', 'CYP4B1', 'VSIG4', 'CDKN2A', 'EBPL', 'CMA1', 'KDR', 'ABCC8', 'IL4', 'S100P', 'SLC4A1', 'CES1', 'SCGB2A1', 'DUSP2', 'TREM2', 'MMP3', 'PCDH10', 'FSCN1', 'NOTCH2', 'GSTA1', 'WNT2', 'ENAH', 'FASN', 'SFXN1', 'TYROBP', 'VAMP8', 'AREG', 'CDK6', 'ADIPOQ', 'BIRC3', 'REXO4', 'C15orf48', 'IFNGR1', 'SLC25A37', 'SMAD3', 'APOC1', 'CTHRC1', 'FHL1', 'SPI1', 'SCG2', 'GKN2', 'OTOP2', 'CACNG4', 'MS4A6A', 'NDUFA4L2', 'ICOSLG', 'IL4I1', 'TENT5C', 'RUNX1T1', 'TNFRSF25', 'PTGDS', 'ITM2C', 'LAMA2', 'NPM3', 'VEGFA', 'DPP6', 'TMIGD1', 'ENTPD1', 'MTRNR2L11', 'IL22RA2', 'ODF2L', 'CD40', 'RSPO3', 'CTSE', 'MARCO', 'FOS', 'COL8A1', 'ECSCR', 'ADAM17', 'APOA5', 'DCLK1', 'RPS4Y1', 'MYO5B', 'NKG7', 'ASCL1', 'SST', 'CEACAM1', 'INSM1', 'TP73', 'DEPP1', 'CYP3A4', 'TXLNA', 'FGL2', 'COTL1', 'SERPINA1', 'TFPI', 'ICA1', 'HAMP', 'MMRN2', 'CYP2A7', 'CCR3', 'GZMB', 'PIM2', 'NF1', 'IFNL1', 'RHOA', 'COL19A1', 'PMP22', 'CXCL5', 'FOXP3', 'RNF43', 'KLK11', 'DSP', 'ASAH1', 'CCL26', 'FOXC2', 'LGR6', 'EDN1', 'OGN', 'ANKRD28', 'CXCR5', 'GPR34', 'CD34', 'CCPG1', 'ID2', 'HPX', 'CD276', 'AQP2', 'F3', 'CDCA7', 'MS4A7', 'COL5A2', 'CYP2F1', 'DAPK3', 'NXPH1', 'CD2', 'TKT', 'SOX17', 'TRAC', 'GPR183', 'SHANK3', 'PCOLCE2', 'MCEMP1', 'SVIL', 'TK1', 'SLC22A8', 'CA2', 'CDKN2C', 'THAP2', 'S
              ...
               'FCN3', 'CDHR5', 'FASLG', 'HSP90B1', 'ORC6', 'NPC2', 'HDC', 'LRRC15', 'IGF1R', 'CTTN', 'SEC61B', 'EPOR', 'CDH1', 'LILRB4', 'NOTCH3', 'ANXA3', 'APCDD1', 'ID4', 'TRDN', 'SLC7A11', 'EIF4EBP1', 'MYBPC1', 'MYOM1', 'IFIT3', 'DUSP5', 'SLC26A2', 'GDF15', 'CDX1', 'FFAR4', 'FOXI1', 'GUCA2B', 'SEC24A', 'SLC29A4', 'GZMK', 'ITGAM', 'PCNA', 'LILRA5', 'EDNRB', 'CDK15', 'EHF', 'PTPRB', 'BAALC', 'SCG5', 'CLEC10A', 'POSTN', 'FGFBP1', 'UPK1B', 'ELOVL5', 'PDK4', 'EPHA2', 'HOXD9', 'GRB14', 'TRAT1', 'PGR', 'SDC1', 'CX3CR1', 'ABCC11', 'GATM', 'TP53', 'HRCT1', 'BEST4', 'XCR1', 'ATM', 'TAGLN', 'CEBPB', 'KRT20', 'SOD3', 'NXPE4', 'CLEC14A', 'FZD7', 'LAG3', 'GNLY', 'C1QB', 'FCN2', 'CP', 'NID1', 'LY6E', 'EMP3', 'STAT2', 'ACTG2', 'CD247', 'ICAM1', 'CSF2RA', 'RGS5', 'PPP1R1B', 'VEGFC', 'CCL21', 'TMEM61', 'COCH', 'S100A4', 'STAT4', 'ARPC3', 'CDC42EP1', 'PIM1', 'IL12RB2', 'MEIS2', 'SDC4', 'MFAP5', 'FEZ1', 'RTN4', 'SLC12A2', 'S100B', 'PRF1', 'SSR2', 'JCHAIN', 'BAMBI', 'ACE2', 'NOSTRIN', 'PELI1', 'PLK1', 'ACACB', 'TFAP2A', 'CCL15', 'FCGR2A', 'GREM1', 'C20orf85', 'S100A1', 'GPC3', 'CSTA', 'FCER1A', 'OSTC', 'MYH14', 'C2orf88', 'SMIM14', 'TPSG1', 'IL12B', 'MYO6', 'C11orf96', 'NT5E', 'CILP', 'RARRES1', 'SFTPD', 'TRBC1', 'CD3D', 'WARS', 'GZMH', 'TUBA1A', 'ETV5', 'TCEAL7', 'TGFBR1', 'NOD1', 'VTN', 'CEACAM7', 'GPX3', 'TRAPPC3', 'C7', 'CTSL', 'GEM', 'PTTG1', 'MRC1', 'APOBEC3A', 'PBK', 'SEC62', 'ID1', 'IGFBP3', 'LDHB', 'MMP11', 'EPHA4', 'CRYBA2', 'PABPC1', 'IL15RA', 'NAT8', 'CFB', 'CD24', 'CLCA2', 'DNASE1L3', 'RGS16'}) and 3 missing columns ({'y_coord', 'cell_id', 'x_coord'}).
              
              This happened while the csv dataset builder was generating data using
              
              hf://datasets/GHISTPlus/GHIST-Plus-bundle/evaluation_data/atera/cell_gene_matrix_filtered.csv (at revision 0932796efd6a92fcc9c4bfeaf447a89ccce51269), ['hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/atera/cell_coords.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/atera/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_coords.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast2/cell_type_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast5k/cell_coords_histology_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/breast5k/cell_gene_matrix_filtered.csv', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/imputation/figure3_coordinates.csv.gz', 'hf://datasets/GHISTPlus/GHIST-Plus-bundle@0932796efd6a92fcc9c4bfeaf447a89ccce51269/evaluation_data/imputation/figure3_ground_truth.csv.gz']
              
              Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

cell_id
int64
x_coord
float64
y_coord
float64
1
77,496.841733
4,082.609656
2
77,537.900401
4,370.606505
3
77,184.576033
3,873.576296
4
77,250.12938
3,869.166524
5
77,384.880698
4,028.793313
6
77,402.927754
4,039.828016
7
77,608.767233
4,698.044269
8
77,262.205517
4,773.734259
10
77,645.965348
4,964.74568
11
77,504.164625
4,436.79834
12
77,533.368381
4,458.082255
13
77,377.910075
4,424.744235
14
77,405.112541
4,455.038862
15
77,491.16299
4,457.368883
16
77,757.498847
4,406.618391
17
77,681.642409
4,687.665622
18
77,767.367421
4,676.995191
19
77,661.391378
4,519.183994
20
77,675.221637
4,090.759088
21
77,707.766352
4,350.739595
22
77,670.613108
4,131.705586
23
75,576.460471
5,207.854871
24
72,702.227606
4,879.310757
25
71,615.779445
4,955.108398
27
71,881.301518
4,835.285078
28
73,490.032672
4,964.085858
29
72,101.87823
4,718.409513
30
74,586.200639
4,941.230864
31
73,822.263766
4,986.84304
32
73,926.958675
5,022.137536
33
74,305.788181
4,863.563015
34
74,314.140861
4,908.361581
36
72,131.051938
5,246.706835
37
74,577.292972
5,265.406334
38
72,293.407026
5,231.636427
39
73,952.353312
5,220.429419
40
73,816.456741
5,189.144562
41
74,731.157576
5,226.575805
42
74,768.826378
5,199.801215
43
74,902.167273
5,206.50993
44
75,432.08101
4,879.138671
45
75,269.804428
4,755.513862
47
76,934.155068
4,929.299145
48
76,532.563511
4,782.766704
49
76,625.775811
4,705.831506
50
76,925.040744
5,072.876723
52
75,675.618704
4,985.018186
53
75,700.048299
4,955.589336
54
75,679.760741
4,838.358601
55
76,360.769952
4,974.524822
56
76,035.965887
5,090.526428
57
76,264.704636
5,153.247145
58
75,707.956881
4,638.28954
60
75,787.707225
4,672.027241
63
76,622.848298
4,638.162573
64
77,227.710211
5,124.79241
65
77,017.890736
5,189.001011
66
77,140.003196
5,199.766397
67
76,903.310249
5,122.009091
68
77,547.153788
5,146.891362
69
77,254.809826
5,044.80904
70
77,094.931419
5,000.251794
71
77,099.166428
4,726.981833
72
77,181.046944
4,752.219104
73
77,288.648005
4,422.325687
74
77,183.830322
4,548.79255
75
77,316.311911
4,502.802414
76
75,746.846489
4,437.075111
77
74,481.789582
4,603.64529
78
74,740.094629
4,578.377749
79
74,614.797527
4,712.509595
81
75,326.258132
4,726.243822
82
75,203.259187
4,636.961879
84
75,432.077192
4,572.807333
85
75,303.960476
4,448.932645
88
73,385.801245
4,437.450125
89
72,769.643886
4,782.034326
90
73,117.966729
4,829.269108
91
73,161.121075
4,594.905998
92
73,166.042679
4,365.272054
93
75,017.656489
4,705.778952
94
74,830.960171
4,440.135673
95
74,883.415466
4,745.827906
97
76,810.227061
4,366.586414
98
76,753.596723
3,984.775452
100
75,644.886195
4,117.98209
101
74,480.035718
4,077.647946
106
75,086.049028
4,287.121509
107
75,865.171736
4,371.789775
110
76,368.870049
4,087.135858
113
72,251.736973
6,483.371723
114
72,210.426341
6,494.434017
115
72,225.502479
6,504.196605
116
72,223.740039
6,471.897379
117
72,175.801668
6,584.235427
118
71,974.014053
6,610.985756
119
71,974.541898
6,591.541713
120
72,174.0108
6,459.324372
121
71,999.141507
6,494.616251
122
72,190.200386
6,505.014207
End of preview.

GHIST+ data and model bundle

This bundle contains the model artifacts, predictions, evaluation inputs, comparison outputs, and plot-ready tables released with GHIST+. Source code is provided in that repository.

Download

hf download GHISTPlus/GHIST-Plus-bundle \
  --repo-type dataset \
  --local-dir bundle

The bundle is approximately 38 GB. Individual files can also be downloaded from this page.

Model checkpoints

The four released GHIST+ checkpoint files are:

  • GHIST_plus/models/breast_multi/ghist_plus_breast_multi_checkpoint.pth
  • GHIST_plus/models/breast_single/ghist_plus_breast_single_checkpoint.pth
  • GHIST_plus/models/imputation/ghist_plus_gene_imputation_checkpoint.pth
  • GHIST_plus/models/pancancer/ghist_plus_pancancer_checkpoint.pth

They exclude the frozen third-party UNI2-H encoder weights. Users must obtain UNI2-H directly from MahmoodLab/UNI2-h, accept its terms, and follow the Pretrained Checkpoints instructions in the GHIST+ README. That procedure reconstructs the complete checkpoints locally at the same filenames, so the existing inference commands and configs remain unchanged.

The bundle does not redistribute UNI2-H weights.

Use with the figure notebooks

The simplest layout is:

download-parent/
├── GHIST_plus/
└── bundle/

Start Jupyter from the code repository root. Figure2.ipynb through Figure5.ipynb will find ../bundle automatically:

cd GHIST_plus
jupyter lab

For another location, set:

export GHIST_BUNDLE_ROOT="/path/to/bundle"
jupyter lab

GHIST_BUNDLE_ROOT must point to the bundle directory, not its parent.

Contents

  • Figure 2: evaluation data, GHIST+ predictions, and comparison-model predictions under evaluation_data/, GHIST_plus/predictions/, and other_models/.

  • Figure 3: imputation inputs and predictions, including the bundled VQ/composition ablation predictions.

  • Figure 4: plot-ready tables under figure_data/figure4/.

  • Figure 5: paired PCC and coverage tables under figure_data/figure5/.

    bundle/ ├── GHIST_plus/ │ ├── models/ │ └── predictions/ ├── evaluation_data/ ├── figure_data/ └── other_models/

The paths and lightweight input schemas were checked against the released Figure 2–5 notebooks.

Data and third-party terms

The bundle combines author-generated artifacts, processed public source data, and outputs from comparison methods. It therefore has no single blanket license; the Hugging Face license field is other. See THIRD_PARTY_NOTICES.md for component-specific sources, versions, attributions, and terms. Upstream terms remain applicable.

Tutorial data

tutorial.ipynb does not use this bundle as its DATA_ROOT. The tutorial requires a separately prepared GHIST data directory containing aligned H&E images, segmentation masks, nuclei metadata, and inputs for the selected training mode.

Downloads last month
32