Datasets:
sample_id stringlengths 8 11 | source_id stringlengths 1 4 | category stringclasses 1
value | url stringlengths 5 47 | title stringlengths 1 199 | scam_type stringclasses 0
values | language stringclasses 0
values | description stringclasses 0
values | annotations_json stringclasses 0
values | image imagewidth (px) 1.88k 59.8k ⌀ | image_path stringlengths 30 33 | image_available bool 2
classes | mhtml_path stringlengths 35 38 | mhtml_bytes unknown | manual_annotation_path stringclasses 0
values | manual_annotation_bytes unknown |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
benign/1 | 1 | benign | https://www.gov.bz | Server Unavailable | null | null | null | null | dataset_benign/截图_benign/1.png | true | dataset_benign/mhtml_benign/1.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93ZWIuYXJjaGl2ZS5vcmc(...TRUNCATED) | null | null | |
benign/10 | 10 | benign | https://www.fiji.gov.fj | Fiji Government - Home | null | null | null | null | dataset_benign/截图_benign/10.png | true | dataset_benign/mhtml_benign/10.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93d3cuZmlqaS5nb3YuZmo(...TRUNCATED) | null | null | |
benign/100 | 100 | benign | masienda.com | Masienda | Heirloom Corn, Masa & Mexican Kitchen Essentials | null | null | null | null | dataset_benign/截图_benign/100.png | true | dataset_benign/mhtml_benign/100.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly9tYXNpZW5kYS5jb20vDQp(...TRUNCATED) | null | null | |
benign/1000 | 1000 | benign | https://www.addmotor.com | Addmotor Electric Bike & Electric Trike Shop- Best E-Bikes | Electric Tricycles For Adults | null | null | null | null | dataset_benign/截图_benign/1000.png | true | dataset_benign/mhtml_benign/1000.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93d3cuYWRkbW90b3IuY29(...TRUNCATED) | null | null | |
benign/1001 | 1001 | benign | https://www.addshopfitting.com | AddShopFitting – Astrid Display & Decor Shop Fitting | null | null | null | null | dataset_benign/截图_benign/1001.png | true | dataset_benign/mhtml_benign/1001.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93d3cuYWRkc2hvcGZpdHR(...TRUNCATED) | null | null | |
benign/1002 | 1002 | benign | https://www.addthis.com | fw_error_www | null | null | null | null | dataset_benign/截图_benign/1002.png | true | dataset_benign/mhtml_benign/1002.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93ZWIuYXJjaGl2ZS5vcmc(...TRUNCATED) | null | null | |
benign/1003 | 1003 | benign | https://www.adecco.ch | Adecco Schweiz - Ihr Partner für Jobs und Rekrutierung | null | null | null | null | dataset_benign/截图_benign/1003.png | true | dataset_benign/mhtml_benign/1003.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93d3cuYWRlY2NvLmNvbS9(...TRUNCATED) | null | null | |
benign/1004 | 1004 | benign | https://www.adecco.gr | Επίσημος ιστότοπος Adecco Greece | null | null | null | null | dataset_benign/截图_benign/1004.png | true | dataset_benign/mhtml_benign/1004.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93d3cuYWRlY2NvLmNvbS9(...TRUNCATED) | null | null | |
benign/1005 | 1005 | benign | https://www.adeje.es | Ayuntamiento de Adeje | null | null | null | null | dataset_benign/截图_benign/1005.png | true | dataset_benign/mhtml_benign/1005.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93d3cuYWRlamUuZXMvaW5(...TRUNCATED) | null | null | |
benign/1006 | 1006 | benign | https://www.adelaidehillswine.com.au | Adelaide Hills Wine | Home | South Australia. | null | null | null | null | dataset_benign/截图_benign/1006.png | true | dataset_benign/mhtml_benign/1006.mhtml | "RnJvbTogPFNhdmVkIGJ5IEJsaW5rPg0KU25hcHNob3QtQ29udGVudC1Mb2NhdGlvbjogaHR0cHM6Ly93d3cuYWRlbGFpZGVoaWx(...TRUNCATED) | null | null |
ScamWeb
A Multimodal Benchmark for Grounded Understanding of Cyber-enabled Fraud Webpages
Code (anonymous mirror) · Dataset · License
Dataset Summary
ScamWeb is a multimodal and multilingual benchmark for understanding cyber-enabled fraud webpages. It combines URL and title metadata, full-page screenshots, archived MHTML evidence, and expert annotations to evaluate whether a system can recognize fraud, identify its subtype, localize deceptive evidence, and explain its decisions.
This repository provides the 5,658-page frozen benchmark, comprising 3,416 expert-annotated scam webpages and 2,242 benign webpages, together with supporting annotation artifacts. The broader ScamWeb collection contains 10,026 webpages; its 4,368 auxiliary machine-annotated scam pages are outside the frozen evaluation corpus.
| Benchmark characteristic | Coverage |
|---|---|
| Frozen benchmark webpages | 5,658 |
| Expert-annotated scam webpages | 3,416 |
| Benign webpages | 2,242 |
| Benchmark scam taxonomy | 25 subtypes |
| Expert evidence regions | 5,489 |
| Language-label combinations in expert annotations | 47 |
Language labels include both individual languages and multilingual combinations. The taxonomy and evaluation protocol are maintained in the accompanying code repository.
Tasks and Evaluation
The four linked tasks form GroundScam, the paper's unified evaluation setting.
| Task | Expected output | Evaluation metrics |
|---|---|---|
| Fraud detection | scam or benign |
Accuracy, precision, recall, F1 |
| Subtype classification | One of 25 scam subtypes | Accuracy, macro precision, macro recall, macro F1 |
| Evidence grounding | Screenshot-space evidence regions | mIoU, P/R/F1@0.5, page Recall@0.3/0.5 |
| Rationale generation | An explanation for each evidence region | ROUGE-L, multilingual BERTScore, grounded coverage |
The accompanying code implements ScamMark, a training-free LLM–VLM baseline that integrates URL, textual, visual, and structural evidence through staged reasoning. For supervised evaluation, the fixed 70/15/15 training/validation/test partition is maintained in the code repository. The Hugging Face configurations expose the complete corpus through the full split.
The formal evaluation retains all 5,658 benchmark records and uses trusted MHTML-derived evidence for 5,482 pages under the code repository's evidence policy. Archived MHTML payloads in this dataset and MHTML content admitted as model input serve distinct roles; follow that policy when reproducing the benchmark.
Dataset Construction
The broader collection was assembled from public anti-scam platforms, scam directories, public reports, and existing datasets between November 2025 and February 2026. Live webpage capture and historical recovery through the Internet Archive preserve visual and structural evidence independently of subsequent website availability. The benign set was drawn from LegitPhish (Potpelwar, Kulkarni, and Waghmare, 2025).
For the expert-annotated scam set, two domain experts independently annotated each webpage, with disagreements adjudicated by a third expert. Annotations include the scam subtype, language, page description, and evidence regions. Each region is represented by a bounding box in original screenshot coordinates (xyxy) and a human-written explanation. Benign webpages provide the input modalities and binary category without scam evidence annotations.
Data Organization
The release uses Zstd-compressed, self-contained Parquet files: 79 shards, totaling 29728933110 bytes (approximately 29.73 GB). Original media and annotation payloads are stored as bytes.
| Configuration | Split | Rows | Contents |
|---|---|---|---|
samples (default) |
full |
5,658 | Benchmark webpage metadata, screenshots, MHTML, and expert annotations |
artifacts |
full |
7,845 | Supporting machine-generated annotations, collection workbook, annotation documentation, and ancillary files |
samples fields
| Field | Representation and meaning |
|---|---|
sample_id |
Unique string identifier with a scam/ or benign/ prefix |
source_id |
Source identifier as a string, including any suffix |
category |
Binary category: scam or benign |
url, title |
Webpage URL and title |
scam_type, language, description |
Expert annotation fields for scam pages; null for benign pages |
annotations_json |
JSON-encoded expert evidence-region annotations; null for benign pages |
image |
Hugging Face Image(decode=False): original PNG bytes and relative path, or None |
image_path, image_available |
Expected relative screenshot path and availability flag |
mhtml_path, mhtml_bytes |
Relative archive path and original MHTML bytes |
manual_annotation_path, manual_annotation_bytes |
Relative expert annotation path and original JSON bytes; null for benign pages |
Expert annotation strings are preserved in their source form. Use the accompanying code's benchmark taxonomy for canonical subtype evaluation. Screenshots are returned as bytes to support selective decoding across varying image resolutions; check image_available before accessing image.
artifacts fields
| Field | Meaning |
|---|---|
path |
Relative artifact path |
kind |
Artifact role, such as machine annotation, workbook, or annotation documentation |
size |
Payload size in bytes |
sha256 |
SHA-256 digest of the payload |
bytes |
Original file bytes |
Machine-generated annotations are auxiliary material and are not expert gold labels. Their presence does not define additional complete benchmark records. For GroundScam evaluation, use the expert annotation fields in samples and the accompanying evaluation protocol.
Usage
Install the Hugging Face Datasets library:
pip install datasets huggingface_hub
Stream benchmark records without downloading the entire release:
import json
from datasets import Image, load_dataset
pages = load_dataset(
"kirco567/ScamWeb-Dataset",
"samples",
split="full",
streaming=True,
token=False,
)
pages = pages.cast_column("image", Image(decode=False))
page = next(iter(pages))
print(page["sample_id"], page["category"])
png_bytes = page["image"]["bytes"] if page["image"] is not None else None
mhtml_bytes = page["mhtml_bytes"]
regions = json.loads(page["annotations_json"]) if page["annotations_json"] else []
Read supporting artifacts separately:
from datasets import load_dataset
artifacts = load_dataset(
"kirco567/ScamWeb-Dataset",
"artifacts",
split="full",
streaming=True,
token=False,
)
artifact = next(iter(artifacts))
print(artifact["path"], artifact["kind"], artifact["size"])
Download the complete release:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="kirco567/ScamWeb-Dataset",
repo_type="dataset",
local_dir="./ScamWeb-Dataset",
token=False,
)
Pin a repository revision when reporting experimental results. The release includes file and shard manifests with SHA-256 digests for integrity verification.
Considerations for Use
ScamWeb supports research on fraud detection, multimodal understanding, grounded explanations, and robust evaluation. Its long-tailed subtype and language distributions should be considered when interpreting aggregate scores. Archived labels describe the collected evidence rather than the current state of a live domain; practical decisions about websites should include human review.
Archived scam pages may contain deceptive, explicit, or otherwise sensitive material. Process webpage archives in an isolated environment and avoid executing embedded content or automatically visiting collected URLs. These handling recommendations do not add restrictions to the license below.
Citation and Related Work
The accompanying paper is an anonymous AAAI 2027 submission:
ScamWeb: A Multimodal Benchmark for Grounded Understanding of Cyber-enabled Fraud Webpages. Anonymous submission, AAAI 2027.
An archival citation will be added when public author and publication information becomes available. The anonymous code mirror linked above provides the benchmark protocol and ScamMark implementation.
Benign webpage source: Potpelwar, R. S.; Kulkarni, U.; and Waghmare, J. 2025. LegitPhish: A large-scale annotated dataset for URL-based phishing detection. Data in Brief, 63: 111972.
License
The contributors' original annotations, dataset documentation, and dataset curation are released under Creative Commons Attribution 4.0 International (CC BY 4.0), to the extent that the contributors hold the rights needed to license those materials. The complete, unmodified license text is provided in LICENSE.
Third-party webpage content, including material reproduced in screenshots and MHTML archives, remains subject to its original rights and applicable terms; this release does not relicense that content. The accompanying source code is governed by its own license.
- Downloads last month
- 24