Text Classification
Transformers
Safetensors
roberta
devign
defect detection
code
Eval Results (legacy)
text-embeddings-inference
Instructions to use claudios/VulBERTa-MLP-MVD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use claudios/VulBERTa-MLP-MVD with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="claudios/VulBERTa-MLP-MVD")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("claudios/VulBERTa-MLP-MVD") model = AutoModelForSequenceClassification.from_pretrained("claudios/VulBERTa-MLP-MVD", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download tokenization_vulberta.py from claudios/VulBERTa-MLP-MVD: direct link, hf CLI and curl.
- Browser
- Download file 1.39 kB
-
https://huggingface.co/claudios/VulBERTa-MLP-MVD/resolve/main/tokenization_vulberta.py
- Command line
-
hf download hf://claudios/VulBERTa-MLP-MVD/tokenization_vulberta.py
-
curl -L -o tokenization_vulberta.py https://huggingface.co/claudios/VulBERTa-MLP-MVD/resolve/main/tokenization_vulberta.py
1.39 kB
| from typing import List | |
| from tokenizers import NormalizedString, PreTokenizedString | |
| from tokenizers.pre_tokenizers import PreTokenizer | |
| from transformers import PreTrainedTokenizerFast | |
| try: | |
| from clang import cindex | |
| except ModuleNotFoundError as e: | |
| raise ModuleNotFoundError( | |
| "VulBERTa Clang tokenizer requires `libclang`. Please install it via `pip install libclang`.", | |
| ) from e | |
| class ClangPreTokenizer: | |
| cidx = cindex.Index.create() | |
| def clang_split( | |
| self, | |
| i: int, | |
| normalized_string: NormalizedString, | |
| ) -> List[NormalizedString]: | |
| tok = [] | |
| tu = self.cidx.parse( | |
| "tmp.c", | |
| args=[""], | |
| unsaved_files=[("tmp.c", str(normalized_string.original))], | |
| options=0, | |
| ) | |
| for t in tu.get_tokens(extent=tu.cursor.extent): | |
| spelling = t.spelling.strip() | |
| if spelling == "": | |
| continue | |
| tok.append(NormalizedString(spelling)) | |
| return tok | |
| def pre_tokenize(self, pretok: PreTokenizedString): | |
| pretok.split(self.clang_split) | |
| class VulBERTaTokenizer(PreTrainedTokenizerFast): | |
| def __init__( | |
| self, | |
| *args, | |
| **kwargs, | |
| ): | |
| super().__init__( | |
| *args, | |
| **kwargs, | |
| ) | |
| self._tokenizer.pre_tokenizer = PreTokenizer.custom(ClangPreTokenizer()) | |