Dataset Viewer
Auto-converted to Parquet Duplicate
messages
listlengths
2
2
[ { "role": "user", "content": "How can hunters identify tampering with Apple’s Transparency, Consent, and Control (TCC) DB?" }, { "role": "assistant", "content": "Apple's Transparency, Consent, and Control (TCC) database represents a critical security mechanism controlling application access to s...
[ { "role": "user", "content": "How can hunters leverage cloud billing anomalies to discover cryptojacking campaigns?" }, { "role": "assistant", "content": "Cloud billing anomaly detection represents a sophisticated approach to identifying cryptojacking campaigns through financial telemetry analys...
[ { "role": "user", "content": "How can hunters leverage DNS RPZ to disrupt known malware C2 domains?" }, { "role": "assistant", "content": "DNS Response Policy Zones (RPZ) represent a sophisticated defensive mechanism enabling security teams to proactively disrupt adversarial command-and-control ...
[ { "role": "user", "content": "How can hunters leverage hardware performance counters to spot speculation attacks?" }, { "role": "assistant", "content": "Hardware performance counters (HPCs) provide granular visibility into CPU microarchitectural events, making them invaluable for detecting specu...
[ { "role": "user", "content": "How can hunters leverage OpenTelemetry traces to hunt cross‐service attack paths?" }, { "role": "assistant", "content": "OpenTelemetry traces provide comprehensive observability across distributed systems by capturing end-to-end request flows through microservices a...
[ { "role": "user", "content": "How can hunters leverage process memory dumps to spot reflective DLL injections?" }, { "role": "assistant", "content": "Process memory dump analysis represents a critical methodology for detecting reflective DLL injection techniques, which adversaries employ to exec...
[ { "role": "user", "content": "How can hunters leverage Sysmon 15’s FileDelete events to surface ransomware precursors?" }, { "role": "assistant", "content": "Sysmon Event ID 15 (FileDelete) provides critical visibility into file deletion activities, making it an essential tool for detecting rans...
[ { "role": "user", "content": "How can hunters measure success when using adversary knowledge to drive breach discovery?" }, { "role": "assistant", "content": "Cyber threat hunters can effectively measure success in breach discovery by implementing quantitative metrics aligned with the NIST Cyber...
[ { "role": "user", "content": "How can hunters spot abuse of NTP amplification inside corporate nets?" }, { "role": "assistant", "content": "NTP (Network Time Protocol) amplification attacks exploit the protocol's design to generate significant traffic volume through relatively small queries, mak...
[ { "role": "user", "content": "How can hunters spot malicious BIOS patching using hardware telemetry?" }, { "role": "assistant", "content": "Malicious BIOS patching detection through hardware telemetry requires sophisticated monitoring of low-level system behaviors that deviate from established b...
[ { "role": "user", "content": "How can hunters spot malicious manipulation of Windows Defender tamper‐protection registry?" }, { "role": "assistant", "content": "Windows Defender tamper-protection mechanisms protect critical security configurations through registry controls under `HKLM\\\\\\\\SOF...
[ { "role": "user", "content": "How can hunters uncover DNS‐over‐HTTPS (DoH) exfiltration when it masquerades as legitimate traffic?" }, { "role": "assistant", "content": "DNS-over-HTTPS (DoH) exfiltration detection requires sophisticated hunting methodologies that distinguish malicious activity f...
[ { "role": "user", "content": "How can hunters use email header anomalies to trace spear‐phish infrastructure reuse?" }, { "role": "assistant", "content": "Email header anomaly detection represents a sophisticated methodology for tracing spear-phish infrastructure reuse through systematic analysi...
[ { "role": "user", "content": "How can hunters use PCAP to reveal QUIC‐based malware tunneling?" }, { "role": "assistant", "content": "Network traffic analysis using PCAP (Packet Capture) files provides cybersecurity analysts with granular visibility into network communications, enabling detectio...
[ { "role": "user", "content": "How can hunting outcomes feed into risk‐based vulnerability management more than reactive reports?" }, { "role": "assistant", "content": "Hunting outcomes can significantly enhance risk-based vulnerability management by providing dynamic, intelligence-driven insight...
[ { "role": "user", "content": "How can identity‐centric CTI support zero‐trust segment design?" }, { "role": "assistant", "content": "Identity-centric Cyber Threat Intelligence (CTI) plays a pivotal role in zero-trust architecture by providing the contextual data necessary to dynamically adjust s...
[ { "role": "user", "content": "How can immutable infrastructure paradigms affect necessity of full disk images?" }, { "role": "assistant", "content": "Immutable infrastructure paradigms fundamentally alter the traditional approach to system recovery and incident response by eliminating the need f...
[ { "role": "user", "content": "How can incident managers contain live business‐critical databases without violating SLA commitments?" }, { "role": "assistant", "content": "Containing live business-critical databases during security incidents requires a sophisticated balance between operational co...
[ { "role": "user", "content": "How can incident responders contain misconfigured S3 bucket exposures without deleting data?" }, { "role": "assistant", "content": "Containment of misconfigured Amazon S3 bucket exposures requires a multi-layered approach that preserves data integrity while preventi...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate Attacker Tactics, Techniques, and Procedures (TTPs) Observed Via Process Execution?" }, { "role": "assistant", "content": "Volatile memory collection for incident response requires c...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate Bloodhound and Active Directory Graphing?" }, { "role": "assistant", "content": "# Volatile Memory Collection for Bloodhound and AD Graphing Investigation\\n\\n## Technical Approach ...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate Credential Attacks with Mimikatz, Metasploit, and many others?" }, { "role": "assistant", "content": "Volatile memory collection for investigating credential-based attacks requires c...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate DCSync?" }, { "role": "assistant", "content": "DCSync attacks represent sophisticated Active Directory exploitation techniques (MITRE ATT&CK T1003.006) where adversaries impersonate ...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate Golden Tickets?" }, { "role": "assistant", "content": "Collecting volatile memory artifacts for Golden Ticket investigation requires careful methodology to preserve evidence integrit...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate Kerberoasting?" }, { "role": "assistant", "content": "Volatile memory collection for Kerberoasting investigations requires careful methodology to preserve critical artifacts while ma...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate Kerberos Attacks?" }, { "role": "assistant", "content": "Volatile memory collection for investigating Kerberos attacks requires careful methodology to preserve critical authenticatio...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate NTDS.DIT theft?" }, { "role": "assistant", "content": "Collecting volatile memory artifacts for investigating NTDS.DIT theft requires careful execution to preserve evidence integrity...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate NTLM Attacks?" }, { "role": "assistant", "content": "Volatile memory collection for investigating NTLM attacks requires careful methodology to preserve critical artifacts while maint...
[ { "role": "user", "content": "How can incident responders safely collect volatile memory artifacts to investigate Technical and Architectural Mitigation Techniques?" }, { "role": "assistant", "content": "Volatile memory collection for investigating technical and architectural mitigation techniqu...
[ { "role": "user", "content": "How can intel on TTP drift over time calibrate red‐team scenario planning?" }, { "role": "assistant", "content": "Threat actor Tactics, Techniques, and Procedures (TTP) evolution necessitates dynamic red-team scenario calibration through systematic intelligence inte...
[ { "role": "user", "content": "How can intelligence collection plans shape incident response actions in real time?" }, { "role": "assistant", "content": "Intelligence collection plans serve as dynamic frameworks that significantly influence real-time incident response actions by providing structu...
[ { "role": "user", "content": "How can intelligence on attacker dwell times calibrate log‐retention strategies?" }, { "role": "assistant", "content": "Intelligence on attacker dwell times provides critical temporal context for optimizing log-retention strategies, directly supporting the NIST Cybe...
[ { "role": "user", "content": "How can inverse threat modeling reveal where adversaries are likely to hide new backdoors?" }, { "role": "assistant", "content": "Inverse threat modeling represents a paradigmatic shift from traditional defensive approaches by systematically analyzing adversary beha...
[ { "role": "user", "content": "How can IOC expiration policies be tuned to account for adaptive adversarial infrastructure?" }, { "role": "assistant", "content": "IOC expiration policies must incorporate dynamic threat intelligence principles to counter adaptive adversaries who routinely rotate i...
[ { "role": "user", "content": "How can Jupyter notebook telemetry be used to identify code execution abuse in data‐science clusters?" }, { "role": "assistant", "content": "Jupyter notebook telemetry provides granular visibility into computational activities within data-science clusters, enabling ...
[ { "role": "user", "content": "How can Kerberoasting be identified in real‐time within large‐scale Windows environments?" }, { "role": "assistant", "content": "Kerberoasting detection in large-scale Windows environments requires a multi-layered approach combining behavioral analytics, log monitor...
[ { "role": "user", "content": "How can Kerberoasting be prevented in modern hybrid identity architectures?" }, { "role": "assistant", "content": "Kerberoasting prevention in modern hybrid identity architectures requires multi-layered security controls addressing both on-premises Active Directory ...
[ { "role": "user", "content": "How can Kerberos Attacks be identified in real‐time within large‐scale Windows environments?" }, { "role": "assistant", "content": "Real-time identification of Kerberos attacks in large-scale Windows environments requires comprehensive monitoring aligned with NIST C...
[ { "role": "user", "content": "How can Kerberos Attacks be prevented in modern hybrid identity architectures?" }, { "role": "assistant", "content": "Kerberos attack prevention in modern hybrid identity architectures requires implementing comprehensive security controls across on-premises Active D...
[ { "role": "user", "content": "How can keylogger telemetry be transformed into early‐warning intelligence for BEC defenses?" }, { "role": "assistant", "content": "Keylogger telemetry can serve as a critical data source for transforming raw keystroke capture into actionable early-warning intellige...
[ { "role": "user", "content": "How can keystroke timing analytics indicate presence of automated adversary tooling?" }, { "role": "assistant", "content": "Keystroke timing analytics leverages temporal patterns in human typing behavior to detect automated adversary tooling through sophisticated be...
[ { "role": "user", "content": "How can knowledge graphs aid hypothesis building in hunting more than in reactive response?" }, { "role": "assistant", "content": "Knowledge graphs excel at hypothesis generation for threat hunting through their inherent ability to model complex, multi-dimensional r...
[ { "role": "user", "content": "How can knowledge of attacker financial incentives influence defense strategy?" }, { "role": "assistant", "content": "Understanding attacker financial incentives fundamentally transforms defensive strategies by enabling risk-based prioritization and resource allocat...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during Bloodhound and Active Directory Graphing?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) enable adversaries to conduct reconnaissance and lateral movement within ...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during Credential Attacks with Mimikatz, Metasploit, and many others?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) enable adversaries to execute malicious activities u...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during DCSync?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) enable adversaries to execute malicious activities using legitimate system tools, significantly reducing de...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during Golden Tickets?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) significantly enhance stealth during Golden Ticket attacks by leveraging legitimate system tools to...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during Kerberoasting?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) significantly enhance stealth during Kerberoasting attacks by leveraging legitimate system tools to ...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during Kerberos Attacks?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) significantly enhance stealth during Kerberos attacks by leveraging legitimate system tools that ...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during NTDS.DIT theft?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) enable adversaries to execute malicious activities using legitimate system tools, significantly red...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during NTLM Attacks?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) significantly enhance stealth during NTLM attacks by leveraging legitimate system tools that are inhe...
[ { "role": "user", "content": "How can living‐off‐the‐land binaries (LOLBins) facilitate stealth during Technical and Architectural Mitigation Techniques?" }, { "role": "assistant", "content": "Living-off-the-land binaries (LOLBins) enable adversaries to execute malicious activities using legitim...
[ { "role": "user", "content": "How can long‐tail indicator analysis provide early warning of low‐and‐slow exfiltration tactics?" }, { "role": "assistant", "content": "Long-tail indicator analysis represents a sophisticated cybersecurity methodology for detecting subtle, persistent threats that tr...
[ { "role": "user", "content": "How can machine learning algorithms analyze Amcache patterns to identify zero-day malware before traditional signatures are available?" }, { "role": "assistant", "content": "Machine learning algorithms can analyze Amcache patterns to identify zero-day malware by lev...
[ { "role": "user", "content": "How can machine learning algorithms be applied to prefetch data to identify previously unknown malware execution patterns?" }, { "role": "assistant", "content": "Machine learning algorithms can enhance prefetch data analysis by identifying novel malware execution pa...
[ { "role": "user", "content": "How can machine learning enhance both forensic analysis and threat hunting capabilities for endpoint security?" }, { "role": "assistant", "content": "Machine learning (ML) significantly enhances forensic analysis and threat hunting through advanced pattern recogniti...
[ { "role": "user", "content": "How can machine learning identify abnormal process access to credential stores across diverse endpoint configurations?" }, { "role": "assistant", "content": "Machine learning (ML) approaches for detecting abnormal process access to credential stores across heterogen...
[ { "role": "user", "content": "How can machine learning models be trained on ShimCache data to identify anomalous execution patterns indicative of compromise?" }, { "role": "assistant", "content": "Machine learning models can leverage ShimCache data through supervised and unsupervised learning ap...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of Attacker Tactics, Techniques, and Procedures (TTPs) Observed Via Process Execution?" }, { "role": "assistant", "content": "Machine learning models can effectively recognize behavioral...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of Bloodhound and Active Directory Graphing?" }, { "role": "assistant", "content": "Machine learning models can be effectively trained to recognize Bloodhound and Active Directory graphi...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of Credential Attacks with Mimikatz, Metasploit, and many others?" }, { "role": "assistant", "content": "Machine learning models for detecting credential-based attacks require sophistica...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of DCSync?" }, { "role": "assistant", "content": "Machine learning models can be effectively trained to recognize Domain Controller (DC) Sync behavioral patterns by implementing supervis...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of Golden Tickets?" }, { "role": "assistant", "content": "Machine learning models can be trained to recognize Golden Ticket behavioral patterns through supervised learning approaches tha...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of Kerberoasting?" }, { "role": "assistant", "content": "Machine learning models can effectively identify Kerberoasting attack patterns by analyzing authentication logs and service ticke...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of Kerberos Attacks?" }, { "role": "assistant", "content": "Machine learning models can effectively recognize Kerberos attack patterns through supervised and unsupervised learning approa...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of NTDS.DIT theft?" }, { "role": "assistant", "content": "Training machine learning models to recognize NTDS.DIT (Active Directory Database) theft requires a multi-faceted approach combi...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of NTLM Attacks?" }, { "role": "assistant", "content": "Machine learning models can effectively recognize NTLM attack patterns through supervised learning approaches that analyze network...
[ { "role": "user", "content": "How can machine learning models be trained to recognize the behavioral patterns of Technical and Architectural Mitigation Techniques?" }, { "role": "assistant", "content": "Machine learning models can be effectively trained to recognize behavioral patterns of Techni...
[ { "role": "user", "content": "How can machine learning models identify abnormal RDP session characteristics like unusual duration or data transfer?" }, { "role": "assistant", "content": "Machine learning models can effectively identify abnormal Remote Desktop Protocol (RDP) session characteristi...
[ { "role": "user", "content": "How can machine learning models identify credential stuffing attacks that use previously breached password databases?" }, { "role": "assistant", "content": "Machine learning models can effectively identify credential stuffing attacks leveraging breached password dat...
[ { "role": "user", "content": "How can machine learning models identify lateral movement patterns that span multiple protocols and authentication methods?" }, { "role": "assistant", "content": "Machine learning (ML) models can effectively identify lateral movement patterns across diverse protocol...
[ { "role": "user", "content": "How can machine learning models identify malicious PowerShell commands that use heavy obfuscation or encoding?" }, { "role": "assistant", "content": "Machine learning models can effectively identify malicious PowerShell commands employing heavy obfuscation through m...
[ { "role": "user", "content": "How can machine learning models trained on execution telemetry identify zero-day exploits before patches are available?" }, { "role": "assistant", "content": "Machine learning models trained on execution telemetry can identify zero-day exploits through behavioral an...
[ { "role": "user", "content": "How can machine‐learning clustering highlight code obfuscation patterns signaling novel malware builders?" }, { "role": "assistant", "content": "Machine-learning clustering algorithms excel at identifying subtle pattern similarities within large datasets, making the...
End of preview. Expand in Data Studio

Security and AppSec Corpus

Description

Teaches domain-specific instruction following and code generation for this expert.

Source

  • Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset
  • AlicanKiraz0/Cybersecurity-Dataset-Heimdall-v1.1
  • mlfoundations-dev/stackexchange_security and 11 more

Formatted for the MoE-orchestrator project (https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target: security.

Format

Each record is a JSON object with a messages field formatted for Qwen3.5's native chat template:

{"messages": [
  {"role": "system", "content": "..."},
  {"role": "user", "content": "..."},
  {"role": "assistant", "content": "..."}
]}

The records are consumed via tokenizer.apply_chat_template(). Special tokens (<|im_start|>, <|im_end|>) are added by the template, never embedded in content.

Splits

  • train: 166,885 records
  • val: 3,405 records

Usage

from datasets import load_dataset
ds = load_dataset("michaelowusuntim6/security-qwen35", split="train")
print(ds[0]["messages"])

License

mixed. Upstream sources keep their own licences - see the source list above and docs/DATASET_SOURCES.md in the MoE-orchestrator repository for per-source detail.

Downloads last month
37