# llama.cpp (Bonsai-8B) on Ubuntu24.04

**URL:** https://forum.ficusonline.com/t/topic/554
**Category:** AI (LLM VLM Multimodal)
**Tags:** llm, llamacpp, bonsai-8b, huggingface
**Created:** [2026 年 4 月 13 日午後 12:49 UTC](https://forum.ficusonline.com/t/topic/554 "2026-04-13T12:49:22Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![tk-fuse](https://forum.ficusonline.com/user_avatar/forum.ficusonline.com/tk-fuse/32/255_2.png) [@tk-fuse](https://forum.ficusonline.com/u/tk-fuse)
#### Post date: [2026 年 4 月 13 日午後 12:49 UTC](https://forum.ficusonline.com/t/topic/554/1 "2026-04-13T12:49:23Z")

</div>

# Bonsai-8B (llama.cpp)のインストール

> <https://askubuntu.com/questions/1461564/install-llama-cpp-locally>

llama.cppは多種多様なAIモデルを演算処理するエンジンで、Bonsai-8BなどのAIモデルを読み込んで利用します。

## llama.cppのインストール

ソースからビルドするか、バイナリをダウンロードします。

#### バイナリファイルのダウンロード

> **[Releases · ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp/releases)**
>
> LLM inference in C/C++. Contribute to ggml-org/llama.cpp development by creating an account on GitHub.

```bash
$ wget https://github.com/ggml-org/llama.cpp/releases/download/b8777/llama-b8777-bin-ubuntu-x64.tar.gz

```

展開して下記ディレクトリへ移動・実行ファイルなどを確認

```bash
$ tar -xzf llama-b8775-bin-ubuntu-x64.tar.gz
$ cd llama-b8775
$ ls
LICENSE libggml-cpu-zen4.so llama-gemma3-cli
libggml-base.so libggml-rpc.so llama-gguf-split
libggml-base.so.0 libggml.so llama-imatrix
libggml-base.so.0.9.11 libggml.so.0 llama-llava-cli
libggml-cpu-alderlake.so libggml.so.0.9.11 llama-minicpmv-cli
libggml-cpu-cannonlake.so libllama.so llama-mtmd-cli
libggml-cpu-cascadelake.so libllama.so.0 llama-mtmd-debug
libggml-cpu-cooperlake.so libllama.so.0.0.8775 llama-perplexity
libggml-cpu-haswell.so libmtmd.so llama-quantize
libggml-cpu-icelake.so libmtmd.so.0 llama-qwen2vl-cli
libggml-cpu-ivybridge.so libmtmd.so.0.0.8775 llama-results
libggml-cpu-piledriver.so llama-batched-bench llama-server
libggml-cpu-sandybridge.so llama-bench llama-template-analysis
libggml-cpu-sapphirerapids.so llama-cli llama-tokenize
libggml-cpu-skylakex.so llama-completion llama-tts
libggml-cpu-sse42.so llama-debug-template-parser rpc-server
libggml-cpu-x64.so llama-fit-params

```

huggingfaceからモデルBonsai-8B-ggufを指定・ダウンロードして起動：オプション` -hf`

```bash
$ ./llama-cli -hf prism-ml/Bonsai-8B-gguf

```

> **[prism-ml/Bonsai-8B-gguf · Hugging Face](https://huggingface.co/prism-ml/Bonsai-8B-gguf)**
>
> We’re on a journey to advance and democratize artificial intelligence through open source and open science.

```auto
$ ./llama-cli -hf prism-ml/Bonsai-8B-gguf
load_backend: loaded RPC backend from /home/ubuntu/llama.cpp/llama-b8775/libggml-rpc.so
load_backend: loaded CPU backend from /home/ubuntu/llama.cpp/llama-b8775/libggml-cpu-haswell.so
Downloading Bonsai-8B.gguf ───────────────────────────────────────── 100%

Loading model...  

▄▄ ▄▄
██ ██
██ ██ ▀▀█▄ ███▄███▄ ▀▀█▄ ▄████ ████▄ ████▄
██ ██ ▄█▀██ ██ ██ ██ ▄█▀██ ██ ██ ██ ██ ██
██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                    ██ ██
                                    ▀▀ ▀▀

build : b8775-920b3e78c
model : prism-ml/Bonsai-8B-gguf
modalities : text

available commands:
  /exit or Ctrl+C stop or exit
  /regen regenerate the last response
  /clear clear the chat history
  /read <file> add a text file
  /glob <pattern> add text files using globbing pattern

> hello

Hello! I'm Bonsai, an AI assistant developed by PrismML. How can I help you today?

```

バイナリではとても遅い。ビルドから再検討中。

---

<div class="post-metadata">

### Author: ![tk-fuse](https://forum.ficusonline.com/user_avatar/forum.ficusonline.com/tk-fuse/32/255_2.png) [@tk-fuse](https://forum.ficusonline.com/u/tk-fuse)
#### Post date: [2026 年 4 月 14 日午前 12:57 UTC](https://forum.ficusonline.com/t/topic/554/2 "2026-04-14T00:57:07Z")

</div>

## llama.cppのバイナリについて

llama.cppをdockerコンテナとして利用する場合、以下ハードウェア環境別のDcokerイメージが用意されています。

> 注）現時点では、Intel CPU搭載のHD GPUにチューニングされたOpenVINO対応のllama.cppでBonsai-8Bモデルは動作しません（1ビット量子化モデル未対応）。

> <https://github.com/ggml-org/llama.cpp/blob/master/docs/docker.md>

ホストPCがインテルCPUでビデオカードを搭載していない場合、 **注1）OpenVINO** または **注2）SYCL** を含む条件でビルドしたバイナリを使用することで、AIのパフォーマンスが大幅に改善するものと思われます。

* * *

### llama.cpp+ OpenVINO

> **注1）OpenVINO** は、インテル製ハードウェア（CPU、GPU、NPU）向けに特化して設計された、高性能なAI推論の最適化とデプロイを行うためのオープンソースツールキットです。  
> [GitHub - openvinotoolkit/openvino: OpenVINO™ is an open source toolkit for optimizing and deploying AI inference · GitHub](https://github.com/openvinotoolkit/openvino?tab=readme-ov-file)

> <https://github.com/ggml-org/llama.cpp/blob/master/docs/backend/OPENVINO.md>

* * *

### llama.cpp + SYCL

> **注2）SYCL** は、CPU・GPU・FPGAなどのさまざまなハードウェアアクセラレータ向けにコードを書く際の開発者の生産性を向上させるために設計された、高水準の並列プログラミングモデルです。異種コンピューティングのための単一ソース言語であり、標準C++17をベースとしています。  
> [SYCL - C++ Single-source Heterogeneous Programming for Acceleration Offload](https://www.khronos.org/sycl/)

> <https://github.com/ggml-org/llama.cpp/blob/master/docs/backend/SYCL.md>

---

<div class="post-metadata">

### Author: ![tk-fuse](https://forum.ficusonline.com/user_avatar/forum.ficusonline.com/tk-fuse/32/255_2.png) [@tk-fuse](https://forum.ficusonline.com/u/tk-fuse)
#### Post date: [2026 年 4 月 15 日午前 2:24 UTC](https://forum.ficusonline.com/t/topic/554/3 "2026-04-15T02:24:03Z")

</div>

## Bonsai-8Bデモ

> **[GitHub - PrismML-Eng/Bonsai-demo: Bonsai Demo](https://github.com/prismml-eng/Bonsai-demo?tab=readme-ov-file)**
>
> Bonsai Demo

所感：8-10世代Intel CPU、ストレージNvme対応SDD、メモリ16GBのホストPCで動作するが、回答の精度では、まだまだ発展途上という印象。

下記で提供されているllama.cppのUbuntu向けにプレビルドされたCPU版バイナリとVulkan版バイナリで評価。

> **[Release prism-b8796-e2d6742 · PrismML-Eng/llama.cpp](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b8796-e2d6742)**
>
> Pre-built binaries (PrismML fork with Q1\_0 1-bit quantization support).
> macOS/iOS:
> 
> macOS Apple Silicon (arm64)
> macOS Apple Silicon (arm64, KleidiAI enabled)
> macOS Intel (x64)
> iOS XCFramework
> 
> Linu...

- [Ubuntu x64 (Vulkan)](https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b8796-e2d6742/llama-prism-b8796-e2d6742-bin-ubuntu-vulkan-x64.tar.gz)  
[Prompt: 5.0 t/s | Generation: 4.1 t/s]

- [Ubuntu x64 (CPU)](https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b8796-e2d6742/llama-prism-b8796-e2d6742-bin-ubuntu-x64.tar.gz)  
[Prompt: 16.6 t/s | Generation: 7.8 t/s]
