llama.cpp, running on a PlayStation 5.

PS5LM runs open language models on the console's own CPU. You chat with them in the PS5's browser, typing with the DualSense.

Open source under GPL-3.0. Runs on a PS5 you own, on firmware 13.60 or lower.

Qwen3.5 0.8B on a PS5 Slim, answering at 21 tokens a second.

How it works

Nothing here is emulated. The PS5 has an 8-core AMD Zen 2 CPU and runs a FreeBSD-based system, so llama.cpp builds for it almost unchanged: three build flags, one small library shim and one patch. This is what runs on the console, in order:

  1. relapsethe browser exploit opens the console (firmware 7.00 to 13.60)
  2. loadera payload loader accepts homebrew programs
  3. llama-serverupstream llama.cpp, built with the open ps5-payload-dev SDK
  4. modela GGUF file from Hugging Face, read into memory
  5. browserthe PS5's own browser opens the chat page the console serves

A phone or laptop on the same network can open the same chat page, and the server speaks the OpenAI API, so other apps can use the PS5 as a small model server.

Measured on the console

13 to 22 tok/s

Qwen3.5 0.8B generating text, with the prompt read at 38 to 67 tokens a second.

5 of 16 cores

The CPU threads a homebrew program may use. System services share them too.

about 6 GB

The memory one homebrew program can hold, from the pool the home screen also uses.

4 seconds

To load a 0.5 GB model from the console's SSD.

What the console taught us

  • llama.cpp starts one busy-waiting thread per core. On a PS5 that takes every core homebrew gets, and the console shuts itself down. The PS5 build leaves two cores free.
  • llama.cpp's own web interface renders blank in the PS5 browser, so PS5LM ships a plain chat page.
  • A model mapped from disk gets paged out once the browser opens; reading it into memory keeps generation fast.

The full notes are in docs/CONSOLE.md.

Roadmap

  • donellama.cpp builds for the PS5 and runs small models on its CPU
  • doneChat in the PS5's own browser, with the DualSense and the on-screen keyboard
  • nextThe PS5LM app: one payload from Payload Manager. Pick a model that fits, download it from Hugging Face on the console, chat.
  • laterBigger models, Qwen 3.8 27B first, in a native app that gets more memory than a payload
  • laterThe GPU, through llama.cpp's Vulkan backend on a community Vulkan driver for the PS5

Try it

Today this is a developer setup: you build on a Mac or Linux machine and start the server on the console from there. The one-step app is next.

# on your computer
git clone --recursive https://github.com/cobanov/PS5LM
cd PS5LM
scripts/setup-sdk.sh && source scripts/env.sh
scripts/build-llama.sh

# console jailbroken, ftpsrv and shsrv running, a model in /data/PS5LM/models
PS5_HOST=192.168.1.50 scripts/ps5-chat.sh model.gguf

The README has the details, and docs/MODELS.md lists models that fit.

Questions

Does this run games or need any Sony code?

No. PS5LM runs open-weight language models and nothing else. No games, keys, firmware or Sony files are part of it.

Why the CPU and not the GPU?

The CPU path needs nothing but the open payload SDK, so it came first. The GPU is the next big step, through a community Vulkan driver for the PS5.

How big a model can it run?

A homebrew payload can hold about 6 GB including the context. Models up to 3B parameters have run so far, and 4-bit models up to about 8B should fit. Bigger ones need a native app, which is on the roadmap.

Can it break my console?

It does not write to system files. Pushing the CPU or memory too hard can freeze the console until you hold the power button; the defaults are set to avoid that.