llama.cpp, running on a PlayStation 5.

PS5LM runs open language models on the console's own CPU. You chat with them in the PS5's browser, typing with the DualSense.

One payload, ps5lm.elf: pick a model, it downloads on the console, chat. Open source under GPL-3.0, for a PS5 you own on firmware 13.60 or lower.

Qwen3.5 0.8B on a PS5 Slim, answering at 21 tokens a second.

How it works

Nothing here is emulated. The PS5 has an 8-core AMD Zen 2 CPU and runs a FreeBSD-based system, so llama.cpp builds for it almost unchanged: three build flags, one small library shim and one patch. This is what runs on the console, in order:

  1. relapsethe browser exploit opens the console (firmware 7.00 to 13.60)
  2. loadera payload loader accepts homebrew programs
  3. llama-serverupstream llama.cpp, built with the open ps5-payload-dev SDK
  4. modela GGUF file from Hugging Face, read into memory
  5. browserthe PS5's own browser opens the chat page the console serves

A phone or laptop on the same network can open the same chat page, and the server speaks the OpenAI API, so other apps can use the PS5 as a small model server.

Measured on the console

13 to 21 tok/s

Qwen3.5 0.8B generating text. Qwen3.5 2B does about 9.

5 of 16 cores

The CPU threads a homebrew program may use. System services share them too.

about 6 GB

The memory one homebrew program can hold, from the pool the home screen also uses.

12 MB/s

Downloading a model from Hugging Face on the console, over 16 connections.

What the console taught us

  • llama.cpp starts one busy-waiting thread per core. On a PS5 that takes every core homebrew gets, and the console shuts itself down. The PS5 build leaves two cores free.
  • llama.cpp's own web interface renders blank in the PS5 browser, so PS5LM ships a plain chat page.
  • While the browser is in front, the system pages a background program's memory out to disk, and a model then crawls. Locking it in memory took Qwen3.5 2B from 0.9 to 9 tokens a second; locking 2.7 GB froze the console, so the limit is 1.6 GB for now.
  • The kernel caps a connection's receive buffer at 64 KB, so one download stream to the CDN stops at 0.9 MB/s. Sixteen in parallel get about 12 MB/s.

The full notes are in docs/CONSOLE.md.

Roadmap

  • donellama.cpp builds for the PS5 and runs small models on its CPU
  • doneChat in the PS5's own browser, with the DualSense and the on-screen keyboard
  • doneThe PS5LM app (v0.1): one payload, a model library in the PS5 browser, downloads on the console, chat
  • nextListed in Payload Manager, and more models measured on more consoles
  • laterBigger models, Qwen 3.8 27B first, in a native app that gets more memory than a payload
  • laterThe GPU, through llama.cpp's Vulkan backend on a community Vulkan driver for the PS5

Install

  1. Jailbreak a PS5 you own on firmware 7.00 to 13.60, with a payload loader such as Payload Manager.
  2. In Payload Manager, turn on Multiple Payload Sources in Settings and add the source https://ps5lm.cobanov.dev/payloads.json. Install PS5LM from the list and load it. (Or download ps5lm.elf from the latest release and load it like any payload.)
  3. The PS5 browser opens the library. Press Download on a model, then Run, then Open chat.

From a phone or computer on the same network, the library is at http://<console IP>:8082 and the chat at port 8081, which also speaks the OpenAI API. Building from source is in the README.

Questions

Does this run games or need any Sony code?

No. PS5LM runs open-weight language models and nothing else. No games, keys, firmware or Sony files are part of it.

Why the CPU and not the GPU?

The CPU path needs nothing but the open payload SDK, so it came first. The GPU is the next big step, through a community Vulkan driver for the PS5.

How big a model can it run?

Up to 1.6 GB of model for now, which is Qwen3.5 2B at 4-bit. The model has to stay locked in memory to run at full speed, and that memory is shared with the home screen. Bigger models need a native app, which is on the roadmap.

Can it break my console?

It does not write to system files. Pushing the CPU or memory too hard can freeze the console until you hold the power button; the defaults are set to avoid that.