You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
My Observations and Thoughts on the llama.cpp Project
Author: zhouwg Translator: I am not a native English speaker; this article was translated by GLM-5.2 Original:我对llama.cpp项目的观察与思考
I began to heavily participate in the llama.cpp project around April 2024 (I invested nearly RMB 20,000 to purchase hardware equipment for development and testing, and accumulated close to 11 months of full-time development time comparable to a full-load employee in a large company. Calculated by purchasing power, 20,000 RMB is roughly equivalent to 10,000 or 20,000 USD). Because some words on its official website moved me, at the very beginning I held very high expectations for this project, thinking it was different from others. The experience of nearly 3 years made me deeply understand one thing: talking is always easier than doing, and this is true regardless of country or project.
1. Preference for Contributions from Large/Well-Known Companies
1.1 Great Tolerance for Contributions from Large American Companies Like Qualcomm
The ggml official website writes: "We built ggml in the spirit of play. Contributors are encouraged to try crazy ideas, build wild demos, and push the edge of what's possible."
However, in the llama.cpp project, I think almost every participant can observe that the maintainers have a preference for contributions from large/well-known companies.
A PR for the RKNN hardware acceleration backend, submitted by an ordinary developer/developer without noted background, was closed by one of the maintainers on the grounds that the PR heavily used AI-assisted development. Yet we can see that in llama.cpp, there is a PR whose author claimed it was entirely generated by AI, and this PR was approved and merged by the original author of llama.cpp. I was deeply shocked after seeing this: this is naked double standard, probably because that PR was submitted to the ggml-metal backend, and ggml-metal happens to be the favorite backend of the original author of llama.cpp.
My understanding of this case: if the PR for the RKNN hardware acceleration backend came from Rockchip, I think this PR would very likely be accepted by the maintainers.
2. Preference for the ggml-metal Backend and ggml-rpc Backend
Once I saw a statement in the llama.cpp project saying that metal is a first-class citizen of llama.cpp. At that time I was very shocked: is this something a developer from Europe should say?
As for the ggml-rpc backend, all participants of the llama.cpp project understand the relationship between the author of this backend and the author of llama.cpp (my first account ban was also related to this backend: nepotism, which is common throughout thousands of years of history in China. I never expected that top European developers would be no exception. Anyone who understands Asian/Chinese culture would understand what "conflict of interest avoidance" means — not bringing unnecessary trouble to relatives who hold enormous power. Because I had disagreements with some PRs submitted by the author of ggml-rpc and had arguments with him, combined with my own other mistakes, this led to my main GitHub account being banned for the first time in July 2024).
3. Three GitHub Accounts of the Author of FastRPC-based ggml-hexagon Were Banned
The ggml official website writes: "We built ggml in the spirit of play. Contributors are encouraged to try crazy ideas, build wild demos, and push the edge of what's possible."
Three GitHub accounts (GitHub: zhouwg / jeffzhou-zhouwg / jeff-zhouwg) of the author of FastRPC-based ggml-hexagon were banned. The reason for the first account being banned in July 2025, I can understand; the second account being banned in August 2026, I can also understand and accept (I edited ggml-org/llama.cpp#26227 too many times, which may have triggered mass email notifications and caused trouble for the maintainers); but the third account being banned in August 2026, I completely cannot understand, especially since the third account had absolutely no intention of submitting any PR and was still ruthlessly banned: a Qualcomm technical expert replied to my review using the third GitHub account in ggml-org/llama.cpp#26501, and had no intention of banning.
My third GitHub account being ruthlessly, coldly, and unreasonably banned prompted me to deeply reflect on this top European developer whom I once worshipped and still hold deep respect for his powerful technical capabilities. Many people in Western countries who do not know the truth view Chinese authorities through colored glasses. After seeing the above practices of the original author of llama.cpp (especially banning the third GitHub account of the author of FastRPC-based ggml-hexagon), would they have a different understanding of the word "dictatorship"?
4. Ambiguous Handling of Contributions from Different Countries/Regions
ggml-webgpu comes from an American developer. It is purely a toy backend. In a very rough stage, the first PR of ggml-webgpu was quickly accepted by the original author of llama.cpp.
ggml-et comes from a Taiwanese developer. It can be regarded as a toy backend. The first PR of the ggml-et backend was a super-large PR, did not go through strict review, and was quickly accepted by some maintainers. The original author of the llama.cpp project and some maintainers who do not understand Asian culture may not understand / find it hard to understand: for many developers from China, no matter how good the PR is or how high the technical level is, if the PR author is a Taiwan independence supporter, it is impossible to get 100% support from Chinese developers.
The author of FastRPC-ggml-hexagon is from China. FastRPC-ggml-hexagon is the acceleration backend for Qualcomm Snapdragon chips that is currently widely used worldwide, and it has performance advantages over Qualcomm's official implementation on multiple models. But for some unknown reason (my guess is that the original author of llama.cpp hopes to maintain a good relationship with Qualcomm and hopes that Qualcomm's employees participate more actively in the llama.cpp project — this kind of practice of sacrificing individuals for the so-called "big picture" and abandoning the original intention / proud values proclaimed on the ggml.ai official website is common throughout thousands of years of Chinese history), it still cannot be accepted.
Of course, to reflect their "fairness / equality / justice" and other proud values, they will also accept contributions from ordinary Chinese developers / developers without noted background information.
5. Where Did slaren Go?
slaren is one of the important contributors to llama.cpp, and is the original author of the llama.cpp backend subsystem. Where did he go? Why did he suddenly disappear from the llama.cpp project?
Although slaren supported the competing PR of PR-12326 in PR-12326 (which had no essential difference from PR-12326, just wrapped it with nice C++), and I had disagreements with him, I still hold deep respect for his powerful technical capabilities. I am very curious: after llama.cpp was acquired by Hugging Face, why did he, as an important contributor to llama.cpp, suddenly disappear from this project?
6. Where Did Huawei Go?
At one time I saw ggml-cann and ggml-opencl coexisting in the llama.cpp project. Employees from Huawei and Qualcomm were competing to contribute to llama.cpp. I was very moved: although Huawei and Qualcomm are well-known competitors, they could still coexist in the same open source project; although humans have many differences, we only have one Earth, mutual respect & seeking common ground while reserving differences is so good.
As a contributor to ggml-cann, why did Huawei suddenly disappear from the llama.cpp project?
7. Touching Moments
The maintainer of Intel's ggml-sycl is very friendly (I submitted a PR that was later reverted. After submitting the PR, I assessed that this PR might have problems and closed it myself, but unexpectedly the Intel maintainer re-opened that PR and approved it, which moved me very much). The maintainer of Nvidia's ggml-cuda is very friendly (I submitted a PR related to Nvidia's old graphics card that I later actively closed, and unexpectedly the Nvidia maintainer actively replied to my post and gave important tips). Although I hold deep respect for the powerful technical capabilities of the Qualcomm technical expert, Intel and Nvidia showed values completely different from Qualcomm's in the llama.cpp project (forcibly closing my PR, directly misleading the judgment of the original author of llama.cpp, and ultimately leading to my second and third GitHub accounts being banned. On one hand, Qualcomm people seem to treat llama.cpp as a place to rack up KPIs, submitting large amounts of PRs that are basically all merged, with code constantly changed back and forth; on the other hand, an acceleration backend from an independent developer based on Qualcomm Hexagon SDK and Qualcomm Hexagon operators that has better performance than Qualcomm's official implementation on certain models cannot be accepted by the original author of llama.cpp and Qualcomm. Looking at the claims on the original author's official website, how ironic. Anyone with basic values of fairness and justice would probably feel this is a huge double standard after seeing this).
The original author of ggml-vulkan is very friendly (I made a low-level mistake and submitted a PR that turned out to be inappropriate, but he was still relatively friendly to me). I have always had a good impression of the Netherlands, and the original author of ggml-vulkan was touching once again.
The core maintainer of an open source project with global influence does not necessarily have to be the original author of the project. FFmpeg is the best example.
Although the original author of llama.cpp has very powerful technical capabilities and has made important contributions to humanity, his major misjudgment and mistakes on PR-27642once again fully prove that he is not suitable to steer an open source project with global influence like llama.cpp. His closing of PR-6869 in July 2024 and banning my main GitHub account (also my first GitHub account) in his personal project, I can understand and accept; his closing of PR-12326 in July 2025 and banning my main GitHub account (also my first GitHub account) in ggml-org, I can understand and accept; but his erroneous judgment and extreme behavior on PR-27642 in August 2026 (the Qualcomm technical expert had no intention of banning, but he, to cater to a large American company like Qualcomm, banned my two newly registered GitHub accounts in a short period of time, especially banning my last newly registered account that had absolutely no intention of submitting PRs) once again fully proves that he is only a developer with very strong technical capabilities and strong personal preferences, but does not possess the ability to lead an open source project with global influence (a good leader should be able to accommodate different voices and participants of different backgrounds, rather than making major decisions based solely on personal preferences).
ngxson is an important contributor to llama.cpp and also the original author of the mtmd subsystem. He has an Asian cultural background, studied in France and has a multicultural background. He is fluent in English/Chinese/French, proficient in C++, proficient in AI technology, and familiar with Android development. Based on my experience of participating in the llama.cpp project for nearly 3 years, I believe he holds a more fair and equal value orientation, and his words and actions are relatively consistent. What is especially valuable is that, facing problematic PRs submitted by Qualcomm employees, he can stick to principles and reject them, without compromising because of the submitter's corporate background. In the new stage where llama.cpp has been acquired by Hugging Face, I think he is very suitable to take on the core maintenance role of llama.cpp.
The management of the llama.cpp project needs reform. A Technical Steering Committee (TSC) should be established to avoid the original author of llama.cpp having the final say on major decisions alone. The llama.cpp project has a large number of contributions from Chinese developers. A developer from China should be added as one of the core members of the TSC, with veto power on major decisions such as account banning. After llama.cpp was acquired by Hugging Face, it is no longer the personal project of its original author. The original author's strong personal preferences should be discarded, which may help llama.cpp go further & have more influence. A top developer who created a project with global influence, if he forgets his original intention, then perhaps he is no longer suitable to be the core maintainer of the project. Only by not forgetting the original intention can one go steadily and far.
Replies
TODO — by zhouwg (Aug 30, 2026)
Add original links to increase credibility.
After the Chinese version is completed, translate it into English using GLM-5.2/GLM-5.3, and consider @ mentioning the CEO & CTO of Hugging Face.
I heard that Nvidia intends to acquire Hugging Face. I hope people from Nvidia can see this post, and then re-evaluate the value of llama.cpp or re-evaluate the management of the llama.cpp project: I am just an ordinary developer from China. In the llama.cpp community, I can be considered quite rule-abiding (except that because my English is not very good, I edited PRs and posts too many times, causing mass email notifications that brought a little trouble to the maintainers). Some maintainers of llama.cpp can ban my account at will; would they dare to do the same to Qualcomm, to American big companies?
Nvidia officially announced the acquisition of Hugging Face. I look forward to it bringing some changes to the llama.cpp project, and I look forward to having one of my GitHub accounts unbanned. It would be best if my GitHub account registered 10 years ago (which is this account) could be unbanned.
Update — by zhouwg (Sep 6, 2026)
The author of llama.cpp is without a doubt a once-in-a-generation genius programmer. A genius programmer can continue to serve as the tech leader of the project, but project management is best handled by a dedicated community manager. The enormous power involving issues such as account banning should not be decided by the author of llama.cpp alone.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
My Observations and Thoughts on the llama.cpp Project
Author: zhouwg
Translator: I am not a native English speaker; this article was translated by GLM-5.2
Original: 我对llama.cpp项目的观察与思考
I began to heavily participate in the llama.cpp project around April 2024 (I invested nearly RMB 20,000 to purchase hardware equipment for development and testing, and accumulated close to 11 months of full-time development time comparable to a full-load employee in a large company. Calculated by purchasing power, 20,000 RMB is roughly equivalent to 10,000 or 20,000 USD). Because some words on its official website moved me, at the very beginning I held very high expectations for this project, thinking it was different from others. The experience of nearly 3 years made me deeply understand one thing: talking is always easier than doing, and this is true regardless of country or project.
1. Preference for Contributions from Large/Well-Known Companies
1.1 Great Tolerance for Contributions from Large American Companies Like Qualcomm
A PR submitted and merged by a Qualcomm technical expert was reverted by the author of llama.cpp. The tolerance and magnanimity between them is remarkable.
Reference: ggml-org/llama.cpp#27433
1.2 Double Standards for Independent Developers
The ggml official website writes: "We built ggml in the spirit of play. Contributors are encouraged to try crazy ideas, build wild demos, and push the edge of what's possible."
However, in the llama.cpp project, I think almost every participant can observe that the maintainers have a preference for contributions from large/well-known companies.
A PR for the RKNN hardware acceleration backend, submitted by an ordinary developer/developer without noted background, was closed by one of the maintainers on the grounds that the PR heavily used AI-assisted development. Yet we can see that in llama.cpp, there is a PR whose author claimed it was entirely generated by AI, and this PR was approved and merged by the original author of llama.cpp. I was deeply shocked after seeing this: this is naked double standard, probably because that PR was submitted to the ggml-metal backend, and ggml-metal happens to be the favorite backend of the original author of llama.cpp.
My understanding of this case: if the PR for the RKNN hardware acceleration backend came from Rockchip, I think this PR would very likely be accepted by the maintainers.
2. Preference for the ggml-metal Backend and ggml-rpc Backend
Once I saw a statement in the llama.cpp project saying that metal is a first-class citizen of llama.cpp. At that time I was very shocked: is this something a developer from Europe should say?
As for the ggml-rpc backend, all participants of the llama.cpp project understand the relationship between the author of this backend and the author of llama.cpp (my first account ban was also related to this backend: nepotism, which is common throughout thousands of years of history in China. I never expected that top European developers would be no exception. Anyone who understands Asian/Chinese culture would understand what "conflict of interest avoidance" means — not bringing unnecessary trouble to relatives who hold enormous power. Because I had disagreements with some PRs submitted by the author of ggml-rpc and had arguments with him, combined with my own other mistakes, this led to my main GitHub account being banned for the first time in July 2024).
3. Three GitHub Accounts of the Author of FastRPC-based ggml-hexagon Were Banned
The ggml official website writes: "We built ggml in the spirit of play. Contributors are encouraged to try crazy ideas, build wild demos, and push the edge of what's possible."
Three GitHub accounts (GitHub:
zhouwg/jeffzhou-zhouwg/jeff-zhouwg) of the author of FastRPC-based ggml-hexagon were banned. The reason for the first account being banned in July 2025, I can understand; the second account being banned in August 2026, I can also understand and accept (I edited ggml-org/llama.cpp#26227 too many times, which may have triggered mass email notifications and caused trouble for the maintainers); but the third account being banned in August 2026, I completely cannot understand, especially since the third account had absolutely no intention of submitting any PR and was still ruthlessly banned: a Qualcomm technical expert replied to my review using the third GitHub account in ggml-org/llama.cpp#26501, and had no intention of banning.My third GitHub account being ruthlessly, coldly, and unreasonably banned prompted me to deeply reflect on this top European developer whom I once worshipped and still hold deep respect for his powerful technical capabilities. Many people in Western countries who do not know the truth view Chinese authorities through colored glasses. After seeing the above practices of the original author of llama.cpp (especially banning the third GitHub account of the author of FastRPC-based ggml-hexagon), would they have a different understanding of the word "dictatorship"?
4. Ambiguous Handling of Contributions from Different Countries/Regions
ggml-webgpu comes from an American developer. It is purely a toy backend. In a very rough stage, the first PR of ggml-webgpu was quickly accepted by the original author of llama.cpp.
ggml-et comes from a Taiwanese developer. It can be regarded as a toy backend. The first PR of the ggml-et backend was a super-large PR, did not go through strict review, and was quickly accepted by some maintainers. The original author of the llama.cpp project and some maintainers who do not understand Asian culture may not understand / find it hard to understand: for many developers from China, no matter how good the PR is or how high the technical level is, if the PR author is a Taiwan independence supporter, it is impossible to get 100% support from Chinese developers.
The author of FastRPC-ggml-hexagon is from China. FastRPC-ggml-hexagon is the acceleration backend for Qualcomm Snapdragon chips that is currently widely used worldwide, and it has performance advantages over Qualcomm's official implementation on multiple models. But for some unknown reason (my guess is that the original author of llama.cpp hopes to maintain a good relationship with Qualcomm and hopes that Qualcomm's employees participate more actively in the llama.cpp project — this kind of practice of sacrificing individuals for the so-called "big picture" and abandoning the original intention / proud values proclaimed on the ggml.ai official website is common throughout thousands of years of Chinese history), it still cannot be accepted.
Of course, to reflect their "fairness / equality / justice" and other proud values, they will also accept contributions from ordinary Chinese developers / developers without noted background information.
5. Where Did slaren Go?
slaren is one of the important contributors to llama.cpp, and is the original author of the llama.cpp backend subsystem. Where did he go? Why did he suddenly disappear from the llama.cpp project?
Although slaren supported the competing PR of PR-12326 in PR-12326 (which had no essential difference from PR-12326, just wrapped it with nice C++), and I had disagreements with him, I still hold deep respect for his powerful technical capabilities. I am very curious: after llama.cpp was acquired by Hugging Face, why did he, as an important contributor to llama.cpp, suddenly disappear from this project?
6. Where Did Huawei Go?
At one time I saw ggml-cann and ggml-opencl coexisting in the llama.cpp project. Employees from Huawei and Qualcomm were competing to contribute to llama.cpp. I was very moved: although Huawei and Qualcomm are well-known competitors, they could still coexist in the same open source project; although humans have many differences, we only have one Earth, mutual respect & seeking common ground while reserving differences is so good.
As a contributor to ggml-cann, why did Huawei suddenly disappear from the llama.cpp project?
7. Touching Moments
The maintainer of Intel's ggml-sycl is very friendly (I submitted a PR that was later reverted. After submitting the PR, I assessed that this PR might have problems and closed it myself, but unexpectedly the Intel maintainer re-opened that PR and approved it, which moved me very much). The maintainer of Nvidia's ggml-cuda is very friendly (I submitted a PR related to Nvidia's old graphics card that I later actively closed, and unexpectedly the Nvidia maintainer actively replied to my post and gave important tips). Although I hold deep respect for the powerful technical capabilities of the Qualcomm technical expert, Intel and Nvidia showed values completely different from Qualcomm's in the llama.cpp project (forcibly closing my PR, directly misleading the judgment of the original author of llama.cpp, and ultimately leading to my second and third GitHub accounts being banned. On one hand, Qualcomm people seem to treat llama.cpp as a place to rack up KPIs, submitting large amounts of PRs that are basically all merged, with code constantly changed back and forth; on the other hand, an acceleration backend from an independent developer based on Qualcomm Hexagon SDK and Qualcomm Hexagon operators that has better performance than Qualcomm's official implementation on certain models cannot be accepted by the original author of llama.cpp and Qualcomm. Looking at the claims on the original author's official website, how ironic. Anyone with basic values of fairness and justice would probably feel this is a huge double standard after seeing this).
The original author of ggml-vulkan is very friendly (I made a low-level mistake and submitted a PR that turned out to be inappropriate, but he was still relatively friendly to me). I have always had a good impression of the Netherlands, and the original author of ggml-vulkan was touching once again.
After my second account was banned, I did not expect that there would be a developer from Germany replying to me, so I registered a third GitHub account specifically to reply (never expected that, just because I did a review on a Qualcomm PR, I was banned again).
8. My Suggestions
The core maintainer of an open source project with global influence does not necessarily have to be the original author of the project. FFmpeg is the best example.
Although the original author of llama.cpp has very powerful technical capabilities and has made important contributions to humanity, his major misjudgment and mistakes on PR-27642 once again fully prove that he is not suitable to steer an open source project with global influence like llama.cpp. His closing of PR-6869 in July 2024 and banning my main GitHub account (also my first GitHub account) in his personal project, I can understand and accept; his closing of PR-12326 in July 2025 and banning my main GitHub account (also my first GitHub account) in ggml-org, I can understand and accept; but his erroneous judgment and extreme behavior on PR-27642 in August 2026 (the Qualcomm technical expert had no intention of banning, but he, to cater to a large American company like Qualcomm, banned my two newly registered GitHub accounts in a short period of time, especially banning my last newly registered account that had absolutely no intention of submitting PRs) once again fully proves that he is only a developer with very strong technical capabilities and strong personal preferences, but does not possess the ability to lead an open source project with global influence (a good leader should be able to accommodate different voices and participants of different backgrounds, rather than making major decisions based solely on personal preferences).
ngxson is an important contributor to llama.cpp and also the original author of the mtmd subsystem. He has an Asian cultural background, studied in France and has a multicultural background. He is fluent in English/Chinese/French, proficient in C++, proficient in AI technology, and familiar with Android development. Based on my experience of participating in the llama.cpp project for nearly 3 years, I believe he holds a more fair and equal value orientation, and his words and actions are relatively consistent. What is especially valuable is that, facing problematic PRs submitted by Qualcomm employees, he can stick to principles and reject them, without compromising because of the submitter's corporate background. In the new stage where llama.cpp has been acquired by Hugging Face, I think he is very suitable to take on the core maintenance role of llama.cpp.
The management of the llama.cpp project needs reform. A Technical Steering Committee (TSC) should be established to avoid the original author of llama.cpp having the final say on major decisions alone. The llama.cpp project has a large number of contributions from Chinese developers. A developer from China should be added as one of the core members of the TSC, with veto power on major decisions such as account banning. After llama.cpp was acquired by Hugging Face, it is no longer the personal project of its original author. The original author's strong personal preferences should be discarded, which may help llama.cpp go further & have more influence. A top developer who created a project with global influence, if he forgets his original intention, then perhaps he is no longer suitable to be the core maintainer of the project. Only by not forgetting the original intention can one go steadily and far.
Replies
TODO — by zhouwg (Aug 30, 2026)
Add original links to increase credibility.
After the Chinese version is completed, translate it into English using GLM-5.2/GLM-5.3, and consider @ mentioning the CEO & CTO of Hugging Face.
I heard that Nvidia intends to acquire Hugging Face. I hope people from Nvidia can see this post, and then re-evaluate the value of llama.cpp or re-evaluate the management of the llama.cpp project: I am just an ordinary developer from China. In the llama.cpp community, I can be considered quite rule-abiding (except that because my English is not very good, I edited PRs and posts too many times, causing mass email notifications that brought a little trouble to the maintainers). Some maintainers of llama.cpp can ban my account at will; would they dare to do the same to Qualcomm, to American big companies?
Update — by zhouwg (Sep 3, 2026)
https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/
Nvidia officially announced the acquisition of Hugging Face. I look forward to it bringing some changes to the llama.cpp project, and I look forward to having one of my GitHub accounts unbanned. It would be best if my GitHub account registered 10 years ago (which is this account) could be unbanned.
Update — by zhouwg (Sep 6, 2026)
The author of llama.cpp is without a doubt a once-in-a-generation genius programmer. A genius programmer can continue to serve as the tech leader of the project, but project management is best handled by a dedicated community manager. The enormous power involving issues such as account banning should not be decided by the author of llama.cpp alone.
All reactions