optimized attention and mlp for performance, add lora monkey patch for models here and GPTQ_For_Llama models using optimization
This commit is contained in:
3
.gitignore
vendored
3
.gitignore
vendored
@@ -5,3 +5,6 @@ llama-13b-4bit
|
||||
llama-13b-4bit.pt
|
||||
text-generation-webui/
|
||||
repository/
|
||||
build/
|
||||
dist/
|
||||
*.egg-info*
|
||||
|
||||
Reference in New Issue
Block a user