A few days ago i could run the inference as normal. Today is just doesn't work anymore.
I downloaded the model and then setup according to the docs.
warning: not compiled with GPU offload support, --gpu-layers option will be ignored
warning: see main README.md for information on enabling GPU BLAS support
build: 3955 (a8ac7072) with cc (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0 for x86_64-linux-gnu
main: llama backend init
main: load the model and apply lora adapter, if any
llama_model_loader: loaded meta data with 24 key-value pairs and 332 tensors from models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf (version GGUF V3 (latest))
llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
llama_model_loader: - kv 0: general.architecture str = bitnet-b1.58
llama_model_loader: - kv 1: general.name str = bitnet2b
llama_model_loader: - kv 2: bitnet-b1.58.vocab_size u32 = 128256
llama_model_loader: - kv 3: bitnet-b1.58.context_length u32 = 4096
llama_model_loader: - kv 4: bitnet-b1.58.embedding_length u32 = 2560
llama_model_loader: - kv 5: bitnet-b1.58.block_count u32 = 30
llama_model_loader: - kv 6: bitnet-b1.58.feed_forward_length u32 = 6912
llama_model_loader: - kv 7: bitnet-b1.58.rope.dimension_count u32 = 128
llama_model_loader: - kv 8: bitnet-b1.58.attention.head_count u32 = 20
llama_model_loader: - kv 9: bitnet-b1.58.attention.head_count_kv u32 = 5
llama_model_loader: - kv 10: tokenizer.ggml.add_bos_token bool = true
llama_model_loader: - kv 11: bitnet-b1.58.attention.layer_norm_rms_epsilon f32 = 0.000010
llama_model_loader: - kv 12: bitnet-b1.58.rope.freq_base f32 = 500000.000000
llama_model_loader: - kv 13: general.file_type u32 = 40
llama_model_loader: - kv 14: tokenizer.ggml.model str = gpt2
llama_model_loader: - kv 15: tokenizer.ggml.tokens arr[str,128256] = ["!", "\"", "#", "$", "%", "&", "'", ...
llama_model_loader: - kv 16: tokenizer.ggml.scores arr[f32,128256] = [0.000000, 0.000000, 0.000000, 0.0000...
llama_model_loader: - kv 17: tokenizer.ggml.token_type arr[i32,128256] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
llama_model_loader: - kv 18: tokenizer.ggml.merges arr[str,280147] = ["Ġ Ġ", "Ġ ĠĠĠ", "ĠĠ ĠĠ", "...
llama_model_loader: - kv 19: tokenizer.ggml.bos_token_id u32 = 128000
llama_model_loader: - kv 20: tokenizer.ggml.eos_token_id u32 = 128001
llama_model_loader: - kv 21: tokenizer.ggml.padding_token_id u32 = 128001
llama_model_loader: - kv 22: tokenizer.chat_template str = {% for message in messages %}{% if lo...
llama_model_loader: - kv 23: general.quantization_version u32 = 2
llama_model_loader: - type f32: 121 tensors
llama_model_loader: - type f16: 1 tensors
llama_model_loader: - type i2_s: 210 tensors
llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'bitnet-b1.58'
llama_load_model_from_file: failed to load model
common_init_from_params: failed to load model 'models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf'
main: error: unable to load model
Error occurred while running command: Command '['build/bin/llama-cli', '-m', 'models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf', '-n', '128', '-t', '2', '-p', 'Hi', '-ngl', '0', '-c', '2048', '--temp', '0.8', '-b', '1', '-cnv']' returned non-zero exit status 1.
A few days ago i could run the inference as normal. Today is just doesn't work anymore.
I downloaded the model and then setup according to the docs.
LOGS: