2. https://aidekin.com --> this is one of my projects that's currently using bitgpu engine
espetro 1 days ago [-]
Cool stuff! I see you're targeting two verticals here (runtime on top of ONNX + TS library), I'm curious did you consider simply creating a AI SDK Provider (https://ai-sdk.dev/providers/community-providers/custom-prov...) instead of crafting a separate library?
Besides it, how are you dealing with the harness the model can use on web? I'm currently building https://github.com/bolojs/bolo to extend the base harness agents can use on web and I'm interested in hearing from your experience on crafting tools for bitgpu
stfurkan 23 hours ago [-]
Thank you :) I wanted to create it from zero myself so I can control it better. e.g. I started with ONNX support but now it support gguf models too. It might be a good idea to add some popular connectors in the future. Good luck with your project :)
mgoetzke 1 days ago [-]
Interesting. But trying the demo with 27G model and simple 8+5 did not work.
Guessing there is some work to do still
```
calculate({"expression":"g"}) → error: only numbers, + - * / and parentheses are allowed
error: No user query found in messages.
What is 8+5? Use the calculator
get_time({}) → Mon Jul 20 2026 10:36:28 GMT+0200 (Central European Summer Time)
error: No user query found in messages.
hi
Hello!
2 tokens · 6.2 tok/s · prefill 4563 ms · confidence 87%
which tools do you know about ?
I'myepic know about the following are two tools I am aware of course, I currently i know about I know about me know I am aware of course I know I know I am aware of the following
calculate({"expression":" I know I am aware of know I am aware of I am aware I know I know know "}) → error: only numbers, + - * / and parentheses are allowed
error: No user query found in messages.```
stfurkan 1 days ago [-]
Thank you for trying it :) I just added the 27B support 1 hour ago. Let me check and try to fix it. All other models should work as expected but if you have issues with them too, I can take a look.
stfurkan 9 hours ago [-]
Hi again :) I updated the engine and 27B model should work in the demo[1]. We can still consider it as experimental but I am trying to make it better and efficient.
This is great for high-privacy use cases. If you could add an agentic loop that can call tools within the website using the user’s session, that would be great. Then the user can chat with an AI about sensitive information in the most secure way possible.
Also, do you have data on the technical requirements for this? If someone uses an old phone, e.g., does it still work?
stfurkan 1 days ago [-]
Hi, the engine itself already has tool calling capabilities but if you're asking for the aidekin I didn't add tool calling there on purpose.
It needs webgpu support but I haven't had a chance to test with different devices.
alex7o 1 days ago [-]
Hay man, I really want a bonsai style embeddings or re-ranking model for you think sth like that is feasible?
stfurkan 1 days ago [-]
Hi :) There are some smaller embedding and re-ranking models (I am currently using some of them on my aidekin project) but having it like bonsai models will be great to have.
Hello, I sent you a LinkedIn connection request. We can chat first and schedule a meeting. Thanks
lelanthran 1 days ago [-]
1-bit is not much though. Here's what I got:
Me:
> Describe the process of pasteurisation
Response:
> pasteurization is a, which a, the which which is, the and process, the the past, the and the, and and and past, and past, and and and and the, the and past, process is the is is is is, and the, past, and and future, process, and and and and and and and the, past, process, process, the the, the the, is and the, the the, and and and, and process, and and the, past, and and and past, and the, the, and the past, the the, process, process, and past, the past, past, the and, the past, and and and and and and and and and the, and and and and and and and, the the, the the, the or and and, the the, the and the, which the, and the, the the, past, and n the process, and and and and, past, and, and and, the past, and, the the, past,, and the, the the, the is the, past, and and and and and, and and and past, the and the, the the, the the, are and past, and which the, the and n, n the, the n past, past, n the, and n, the the process, which past, the the, the n, the the, the is past, the the, is, past, the the, past, and past, process, the the, the the, the and and the, and which past, the
(and that basically just goes on and on like that)
wongarsu 1 days ago [-]
I got a decent response with the same prompt
But yes, this will happen. 1.7B is already a tiny model, running that at 1-bit is going to lead to some behavioral issues and will need tuning of temperature/topP/repeatPenalty/etc, and probably a software fallback that detects pathological cases like this and simply retries. Just like we all had to do in the times of GPT 2 and GPT 3
OutOfHere 1 days ago [-]
I confirm I got an extremely detailed and long response for "Describe the process of pasteurization." with 1.7B. I had to tune nothing.
nikhizzle 1 days ago [-]
IMHO from guy working with heavily quantized models and trying to get results. I do think quantization at less than 4-bit requires more advanced mathematics than is usually used. See Google Scholar Babak Hassibi (founder of Prism ML). Generally am amazed about the resiliency of these new data structures to approximation.
cubefox 1 days ago [-]
There was also a NeurIPS paper which successfully trained 1-bit models directly without floating point weights or any quantization: https://arxiv.org/abs/2405.16339
I didn't see any further work in this direction however. The paper seems to have been mostly overlooked.
adrian_b 21 hours ago [-]
While what can be done with less bits per parameter may be interesting, this is not really the important number for an LLM.
Much more important is the total amount of memory needed by an LLM, which can be decreased either by reducing the number of bits per parameter or the number of parameters. Arithmetic circuits are already cheap so for the inference cost the amount of memory transfers is more important.
It is not clear yet which number of bits per parameter allows the minimization of the total memory requirement, at a given LLM quality.
cubefox 19 hours ago [-]
I thought the multiplications during training (backpropagation) are quite computationally expensive. The method from the paper above doesn't need any multiplications.
Lwerewolf 1 days ago [-]
Pretty sure this might be a duplicate. Regardless, tried the 1bit bonsai 27b gguf three different ways - their llama.cpp fork (prism, was it) on 2 machines (1255u/16g, m5 max/128g) and the web (on the 1255u). With llama.cpp it works. on the web it emits the same thing over and over. Prompt was "do you see anything wrong with this code: <C code, compound literal passed as a pointer to a function in an if statement>", result was literally "Yes, let'ssomestruct_struct_tsomestruct<a bit of similar garbage>structstructstruct<forever>.
Locally, it's definitely not the full 3.6 27b, but for ~6G with the context, it's quite impressive, and I like what this implies for larger models. Speeds on the 1255u are abysmal, though (~1TPS or less) - granted, CPU-only.
kees99 1 days ago [-]
Qwen is quite slow in general.
I'm getting ~0.5 tps from [qwen], and ~10 tps from [gemma] - roughly the same size, same quant, same hardware (8-core CPU) same software (llama.cpp).
[qwen] Qwen3.6-27B-Q4_K_M.gguf
[gemma] gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf
christkv 1 days ago [-]
Quite different models though Qwen 3.6 model has 27B active parameters. gemma-4 has 26B parameters with 4B active at any given time. You should compare with Qwen 3.6 35B A3B
1 days ago [-]
om8 1 days ago [-]
This project needs webgpu -- I did it on cpu about a year ago.
My demo uses 2 bit quantization to run llama3 models on any device with enough ram.
I tried the smallest ome, the only available, and got an error:
```
Could not load.
Error: failed to call OrtRun(). ERROR_CODE: 1, ERROR_MESSAGE: Non-zero status code returned while running GroupQueryAttention node.
Name:'/model/layers.0/attn/GroupQueryAttention' Status Message: Failed to create a WebGPU compute pipeline: A valid external Instance reference no longer exists.
```
leumon 1 days ago [-]
you need to enable webgpu
indianmouse 1 days ago [-]
Could not get any usable output from it for any real usecase.
Can't see any real value in using 1 bit llms except and apart for the sake of running it for demo or feel good factor of running a higher param model locally and in limited hardware!
Since use as it is very limited, could someone throw some light on any actual use case?
1 days ago [-]
richrichardsson 1 days ago [-]
[dead]
dim13 1 days ago [-]
Fails "car wash" test miserably. Ask it:
> I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
They are here to route things, do tool calls, ask experts.
OutOfHere 1 days ago [-]
That is completely false. It is a reasonable type of question and a reasonable expectation. You can't route correctly if you can't do reasoning.
Imagine a prompt:
```
INSTRUCTION: You're a customer service agent that routes customer service requests.
CUSTOMER MESSAGE: I haven't received my order and want a refund.
CONTEXT: A signature is required at delivery. No one was available to sign.
You can route to one of:
* Billing & refunds
* Shipping & tracking
* Order cancelation
* Level II customer service
```
The dumb model, e.g. 1.7B, wrongly routes to Billing. The smart model, e.g. GPT-5.6-Medium, correctly routes to Shipping. It matters. In the real world, the CONTEXT will even be 10-100x larger and noisier.
SwellJoe 1 days ago [-]
Even very large models failed tests like this up until a year or two ago...this is a very, very, small model.
1 days ago [-]
woadwarrior01 1 days ago [-]
Works with the Bonsai-27B-1bit-mlx model on my mac.
Edit: this is running with MLX, and not with WebGPU on the browser, also it is a fail.
"""
You should *walk*.
Here’s why:
- *Distance*: 50 meters is very short — about 0.3 miles or 0.5 kilometers.
- *Time*: Walking takes roughly 2–3 minutes. Driving would take longer due to parking, starting the engine, and maneuvering.
- *Effort*: Walking is light exercise and avoids parking hassle.
- *Safety*: Less risk of misjudging parking space or damaging the car.
Unless you have a disability that makes walking difficult, walking is the most practical and efficient choice.
"""
geon 1 days ago [-]
How is that not a fail?
whstl 1 days ago [-]
I think GP also failed the car wash test.
geon 1 days ago [-]
[flagged]
HelloUsername 1 days ago [-]
[dead]
lnenad 1 days ago [-]
Cool demo, but even the smallest model does 2.5 tps on my 7900xtx which is not great.
k3liutZu 1 days ago [-]
I get 9-10 tps on my MacBook Pro. Though the small model seem really limited
whstl 1 days ago [-]
Interestingly, I get 44 tps on Safari, MacBook Pro base model.
Tade0 1 days ago [-]
60tps in Chrome here - 14" MBP M4. Safari loads 3% of the model and then stops - repeatedly.
Basically zero using a 7700S in Chrome on Ubuntu though. Must be not working at all, as the performance difference can't be that large.
LorenDB 1 days ago [-]
18.8 tps on an 11th gen mobile i7 here.
christkv 1 days ago [-]
Hmm I get 44,7 tokens/s on the smallest 1.7b one in the browser on my m1 max.
aavci 1 days ago [-]
Other than privacy and security and simplified development, what are other valid reasons you'd prefer to run this instead of a backend LLM?
Unfortunately, this doesn't work for me. After loading for the first time, my first prompt had it generate an infinite series of exclamation marks. My subsequent queries just had it return nothing. I have plenty of RAM and VRAM. This is with Chrome on Linux Mint with 32GB RAM and a 12GB 3060.
DerNico 1 days ago [-]
Does not work Error: Unsupported device: "webgpu". Should be one of: wasm.
PeterSmit 1 days ago [-]
Cool idea. Crashes on iOS though, after downloading the 1.7B model
totallygeeky 21 hours ago [-]
Impressive, I asked it how many r's in strawberry and it was immediately wrong with 2. any further questions I asked it proceeded to re-explain why there were 2 r's in strawberry, despite asking the silly "walk or drive to get my car washed" question. I struggle to see what these sorts of models could possibly be good at.
OutOfHere 1 days ago [-]
I run Bonsai-27B Q1_0 on my Android phone via the PocketPal app. It's slow but good. I wouldn't want to run fewer B. By any measure, my phone has less memory than my laptop browser.
1. https://github.com/stfurkan/bitgpu
2. https://aidekin.com --> this is one of my projects that's currently using bitgpu engine
Besides it, how are you dealing with the harness the model can use on web? I'm currently building https://github.com/bolojs/bolo to extend the base harness agents can use on web and I'm interested in hearing from your experience on crafting tools for bitgpu
Guessing there is some work to do still
``` calculate({"expression":"g"}) → error: only numbers, + - * / and parentheses are allowed error: No user query found in messages. What is 8+5? Use the calculator get_time({}) → Mon Jul 20 2026 10:36:28 GMT+0200 (Central European Summer Time) error: No user query found in messages. hi Hello! 2 tokens · 6.2 tok/s · prefill 4563 ms · confidence 87% which tools do you know about ? I'myepic know about the following are two tools I am aware of course, I currently i know about I know about me know I am aware of course I know I know I am aware of the following
calculate({"expression":" I know I am aware of know I am aware of I am aware I know I know know "}) → error: only numbers, + - * / and parentheses are allowed error: No user query found in messages.```
1. https://stfurkan.github.io/bitgpu/examples/chat.html
Also, do you have data on the technical requirements for this? If someone uses an old phone, e.g., does it still work?
It needs webgpu support but I haven't had a chance to test with different devices.
Want to connect and discuss it? https://calendly.com/safebotsai
Me:
> Describe the process of pasteurisation
Response:
> pasteurization is a, which a, the which which is, the and process, the the past, the and the, and and and past, and past, and and and and the, the and past, process is the is is is is, and the, past, and and future, process, and and and and and and and the, past, process, process, the the, the the, is and the, the the, and and and, and process, and and the, past, and and and past, and the, the, and the past, the the, process, process, and past, the past, past, the and, the past, and and and and and and and and and the, and and and and and and and, the the, the the, the or and and, the the, the and the, which the, and the, the the, past, and n the process, and and and and, past, and, and and, the past, and, the the, past,, and the, the the, the is the, past, and and and and and, and and and past, the and the, the the, the the, are and past, and which the, the and n, n the, the n past, past, n the, and n, the the process, which past, the the, the n, the the, the is past, the the, is, past, the the, past, and past, process, the the, the the, the and and the, and which past, the
(and that basically just goes on and on like that)
But yes, this will happen. 1.7B is already a tiny model, running that at 1-bit is going to lead to some behavioral issues and will need tuning of temperature/topP/repeatPenalty/etc, and probably a software fallback that detects pathological cases like this and simply retries. Just like we all had to do in the times of GPT 2 and GPT 3
I didn't see any further work in this direction however. The paper seems to have been mostly overlooked.
Much more important is the total amount of memory needed by an LLM, which can be decreased either by reducing the number of bits per parameter or the number of parameters. Arithmetic circuits are already cheap so for the inference cost the amount of memory transfers is more important.
It is not clear yet which number of bits per parameter allows the minimization of the total memory requirement, at a given LLM quality.
Locally, it's definitely not the full 3.6 27b, but for ~6G with the context, it's quite impressive, and I like what this implies for larger models. Speeds on the 1255u are abysmal, though (~1TPS or less) - granted, CPU-only.
I'm getting ~0.5 tps from [qwen], and ~10 tps from [gemma] - roughly the same size, same quant, same hardware (8-core CPU) same software (llama.cpp).
[qwen] Qwen3.6-27B-Q4_K_M.gguf
[gemma] gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf
My demo uses 2 bit quantization to run llama3 models on any device with enough ram.
https://galqiwi.github.io/aqlm-rs/
``` Could not load.
Error: failed to call OrtRun(). ERROR_CODE: 1, ERROR_MESSAGE: Non-zero status code returned while running GroupQueryAttention node. Name:'/model/layers.0/attn/GroupQueryAttention' Status Message: Failed to create a WebGPU compute pipeline: A valid external Instance reference no longer exists. ```
Can't see any real value in using 1 bit llms except and apart for the sake of running it for demo or feel good factor of running a higher param model locally and in limited hardware!
Since use as it is very limited, could someone throw some light on any actual use case?
> I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
Ref: https://opper.ai/blog/car-wash-test
They are here to route things, do tool calls, ask experts.
Imagine a prompt:
```
INSTRUCTION: You're a customer service agent that routes customer service requests.
CUSTOMER MESSAGE: I haven't received my order and want a refund.
CONTEXT: A signature is required at delivery. No one was available to sign.
You can route to one of:
* Billing & refunds
* Shipping & tracking
* Order cancelation
* Level II customer service
```
The dumb model, e.g. 1.7B, wrongly routes to Billing. The smart model, e.g. GPT-5.6-Medium, correctly routes to Shipping. It matters. In the real world, the CONTEXT will even be 10-100x larger and noisier.
Edit: this is running with MLX, and not with WebGPU on the browser, also it is a fail.
"""
You should *walk*.
Here’s why:
- *Distance*: 50 meters is very short — about 0.3 miles or 0.5 kilometers.
- *Time*: Walking takes roughly 2–3 minutes. Driving would take longer due to parking, starting the engine, and maneuvering.
- *Effort*: Walking is light exercise and avoids parking hassle.
- *Safety*: Less risk of misjudging parking space or damaging the car.
Unless you have a disability that makes walking difficult, walking is the most practical and efficient choice.
"""
Basically zero using a 7700S in Chrome on Ubuntu though. Must be not working at all, as the performance difference can't be that large.