add test result
This commit is contained in:
@@ -9,6 +9,10 @@ pip install git+https://github.com/johnsmith0031/alpaca_lora_4bit@winglian-setup
|
|||||||
|
|
||||||
Better inference performance with text_generation_webui, about <b>40% faster</b>
|
Better inference performance with text_generation_webui, about <b>40% faster</b>
|
||||||
|
|
||||||
|
Simple expriment results:<br>
|
||||||
|
7b model with groupsize=128 no act-order<br>
|
||||||
|
improved from 13 tokens/sec to 20 tokens/sec
|
||||||
|
|
||||||
<b>Step:</b>
|
<b>Step:</b>
|
||||||
1. run model server process
|
1. run model server process
|
||||||
2. run webui process with monkey patch
|
2. run webui process with monkey patch
|
||||||
|
|||||||
Reference in New Issue
Block a user