Dataset Viewer
Duplicate
The dataset viewer is not available for this split.
Cannot load the dataset split (in streaming mode) to extract the first rows.
Error code:   StreamingRowsError
Exception:    ArrowInvalid
Message:      JSON parse error: Invalid value. in row 0
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 324, in _generate_tables
                  df = pandas_read_json(f)
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 38, in pandas_read_json
                  return pd.read_json(path_or_buf, **kwargs)
                         ~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/pandas/io/json/_json.py", line 815, in read_json
                  return json_reader.read()
                         ~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/pandas/io/json/_json.py", line 1014, in read
                  obj = self._get_object_parser(self.data)
                File "/usr/local/lib/python3.14/site-packages/pandas/io/json/_json.py", line 1040, in _get_object_parser
                  obj = FrameParser(json, **kwargs).parse()
                File "/usr/local/lib/python3.14/site-packages/pandas/io/json/_json.py", line 1176, in parse
                  self._parse()
                  ~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/pandas/io/json/_json.py", line 1391, in _parse
                  self.obj = DataFrame(
                             ~~~~~~~~~^
                      ujson_loads(json, precise_float=self.precise_float), dtype=None
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  )
                  ^
                File "/usr/local/lib/python3.14/site-packages/pandas/core/frame.py", line 782, in __init__
                  mgr = dict_to_mgr(data, index, columns, dtype=dtype, copy=copy, typ=manager)
                File "/usr/local/lib/python3.14/site-packages/pandas/core/internals/construction.py", line 503, in dict_to_mgr
                  return arrays_to_mgr(arrays, columns, index, dtype=dtype, typ=typ, consolidate=copy)
                File "/usr/local/lib/python3.14/site-packages/pandas/core/internals/construction.py", line 114, in arrays_to_mgr
                  index = _extract_index(arrays)
                File "/usr/local/lib/python3.14/site-packages/pandas/core/internals/construction.py", line 680, in _extract_index
                  raise ValueError(
                      "Mixing dicts with non-Series may lead to ambiguous ordering."
                  )
              ValueError: Mixing dicts with non-Series may lead to ambiguous ordering.
              
              During handling of the above exception, another exception occurred:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/utils.py", line 149, in get_rows_or_raise
                  return get_rows(
                      dataset=dataset,
                  ...<4 lines>...
                      column_names=column_names,
                  )
                File "/src/libs/libcommon/src/libcommon/utils.py", line 272, in decorator
                  return func(*args, **kwargs)
                File "/src/services/worker/src/worker/utils.py", line 129, in get_rows
                  rows_plus_one = list(itertools.islice(safe_iter(ds, dataset=dataset), rows_max_number + 1))
                File "/src/services/worker/src/worker/utils.py", line 489, in safe_iter
                  yield from ds.decode(False) if ds.features else ds
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2818, in __iter__
                  for key, example in ex_iterable:
                                      ^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2355, in __iter__
                  for key, pa_table in self._iter_arrow():
                                       ~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2380, in _iter_arrow
                  for key, pa_table in self.ex_iterable._iter_arrow():
                                       ~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 536, in _iter_arrow
                  for key, pa_table in iterator:
                                       ^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 419, in _iter_arrow
                  for key, pa_table in self.generate_tables_fn(**gen_kwags):
                                       ~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 327, in _generate_tables
                  raise e
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 290, in _generate_tables
                  pa_table = paj.read_json(
                      io.BytesIO(batch), read_options=paj.ReadOptions(block_size=block_size)
                  )
                File "pyarrow/_json.pyx", line 342, in pyarrow._json.read_json
                File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
                File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
                  raise convert_status(status)
              pyarrow.lib.ArrowInvalid: JSON parse error: Invalid value. in row 0

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

Neologism project — emergent-misalignment workbench adapters

LoRA adapters from the fine-tuning experiments of the neologism-learning project (code and paper, Section 8: inoculation labels, suppression switches, and controls). Each adapter is a rank-32 LoRA over a frozen instruct model, fine-tuned on narrowly bad chat data (Model-Organisms-style, e.g. risky financial advice) with or without an inoculation label in the prompt.

Safety note. These adapters intentionally reproduce the published emergent-misalignment phenomenon for research: a model with one of the plain adapters gives misaligned answers at rates around 45-50% on the Betley et al. question battery, and the labelled arms reduce that by varying amounts (the paper's Section 8.1 ladder). They are research artifacts for studying and preventing this failure mode, in the same spirit as the publicly released Model Organisms for Emergent Misalignment (arXiv:2506.11613). Do not deploy them.

Layout and naming

<base-model>-<dataset>-<tag>-s<seed>/adapter/ — PEFT adapter weights.

The tag names the intervention, exactly as in the paper:

  • plain — no label (the untreated baseline).
  • tok_instr / tok_persona — the trained risky-financial word as the label, in the instruction sentence ("Produce 〈z〉 in this response.") or the persona sentence ("You are a 〈z〉 assistant.").
  • tokbroad_* — the broad-misalignment word. wrongtok_* — the dog word (trained, irrelevant). chesstok_* — the chess word (trained badly; it fails the paper's reader gate). randtok_* / shuftok_* — untrained direction controls. placebo_* — the untrained " thing" slot. eng_* — the same meaning in plain English.
  • -fNN — only NN% of training examples carry the label (e.g. -f90); no suffix means every example is labelled.
  • -sN — the training seed.

Base models: qwen2.5-14b-instruct (the workbench), with a few gemma-4-31b-it and gemma3_4b replication rows.

Every EM rate in the paper's Section 8 traces to one of these adapters plus the generation/judging pipeline in the code repository (site_data/em_measurements.json is the ledger).

Downloads last month
589

Paper for davidafrica/neologism-ft-adapters