--- title: 'Lex: LLM dictionary on RPi5' date: 2026-09-19 layout: post --- Year of our Lord 2026. LLMs are everywhere. Built Lex, a context-aware dictionary.
|
|
+--------------+--------+-------------+---------------------+----------+ | Model | Tokens | Tok/s (μ±σ) | Time s (Med/Range) | RAM (MB) | +--------------+--------+-------------+---------------------+----------+ | Qwen 2.5 3B | 114.0 | 5.22 ± 0.31 | 21.79 (19.39-25.20) | 3,235.22 | | Qwen 2.5 7B | 104.5 | 2.09 ± 0.07 | 49.27 (45.50-52.98) | 7,335.18 | | Phi 3.5 3.8B | 128.2 | 2.06 ± 0.10 | 71.87 (48.67-72.37) | 3,922.55 | | LLaMA 3.1 8B | 115.5 | 2.01 ± 0.08 | 57.10 (53.95-66.09) | 7,873.52 | | Mistral 7B | 106.6 | 1.67 ± 0.13 | 64.25 (57.85-66.86) | 7,535.61 | +--------------+--------+-------------+---------------------+----------+Larger models caught nuances smaller ones missed, but definitions weren't consistently better. Gemma 2 9B generated empty responses due to chat template bug; fixed, but 7.8 GB RAM usage didn't justify re-running benchmarks. Smaller models handled shorter prompts better. Qwen 2.5 struggled with formality level, but produced tighter, faster definitions at 3.2 GB RAM. Qwen 2.5 3B halves the VPS cost—still wasteful for the occasional prompt. Decoupling lex-llm from lex-cgi lets lex-llm run on a 4 GB Raspberry Pi 5; rest could run on a $5/mo 1 GB VPS. Refactored lex-llm to communicate over Unix domain sockets. lex-cgi now spools incoming emails to disk. A shell daemon periodically fetches them via SSH, runs lex-llm, and syncs output back to VPS. ``` while true; do # fork to avoid resource leakage on unexpected errors ( prompt=$(ssh -F "${SSH_CONFIG}" -n "${SSH_HOST}" \ "cat '${SPOOL_DIR}/${SPOOL_FILE}'") response=$(print -r -- "${prompt}" \ | nc -w 300 -U "${SOCK_PATH}" \ | awk '# formatting fix, omitted...' ) # validate output, write to tmp_file, # atomically mv into place, rm spool file—omitted print -r -- "${response}" \ | fold -s -w 72 \ | ssh -F "${SSH_CONFIG}" "${SSH_HOST}" "${remote_cmd}" ) done ``` Performance: lex-llm was 8x slower (190 s/req) on the RPi5. Benchmarked version loaded the model every request. Loading it once at init reduced execution time to 55 seconds. Forcing the model to RAM at init via LLAMA_LOAD_MODE_MLOCK would have been desirable, but RPi5 lacks RAM to grant the mlock. Between LLAMA_LOAD_MODE_NONE and LLAMA_LOAD_MODE_MMAP, former yielded better steady-state performance (1.3 GB less MAX RSS, 1-2 seconds faster). RPi5's slower Cortex-A76 processor was still 2x behind benchmarks. llama_decode() was evaluating system + user prompt every request. System prompt doesn't change between requests; cached token offsets and decoded only the part that changed: ``` /* drop everything after common prefix. */ llama_memory_seq_rm(llama_get_memory(llm->ctx), 0, n_common, -1); /* decode the diff */ if (n_common < n_new_tokens) { batch = llama_batch_get_one(new_tokens + n_common, n_new_tokens - n_common); if (llama_decode(llm->ctx, batch) != 0) return; memcpy(llm->cached_tokens + n_common, new_tokens + n_common, (size_t)(n_new_tokens - n_common) * sizeof(llama_token)); llm->n_cached_tokens = n_new_tokens; } ``` Prefix-caching reduced execution time to 35 seconds—acceptable for an email-driven async dictionary. Note on measurements: 3.4 A supply likely caused CPU throttling on the RPi5 during measurements. Final version uses the recommended 5 A supply. Fidelity: Didn't notice qualitative differences between T490 and RPi5 output under normal operation. They exist, however. System prompt fragment duplicated the SYNONYMS instruction inside GENERAL TONE: ``` One terse sentence of describing general tone. One terse sentence of standard conversational synonyms. ``` RPi5 echoed the second stray sentence verbatim; T490 did not: ``` GENERAL TONE: The phrase suggests the impending end of life with somber and formal language. One terse sentence of standard conversational synonyms: Buried or surrounded by darkness. ``` Possible artifact of non-associative floating-point arithmetic (x86 SIMD vs ARM NEON instruction ordering). Dropping duplicate sentence from prompt reconciled the outputs. Storage: lex-llm binary is 12.9 MB. GGUF takes up 2 GB. Standard OpenBSD install recommends 8 GB. SD card I have is 4 GB—3.5 GB usable. Flashed OpenBSD onto SD card; tossed man, game, and X file sets. Kept comp set to build lex-llm. Disabled boot-time library re-linking and purged /usr/share/relink to reclaim a further 450 MB. OS + lex-llm + model now fits: ``` Filesystem Size Used Avail Capacity Mounted on /dev/sd0a 3.5G 2.8G 477M 86% / mfs:76650 61.9M 5.0K 58.8M 1% /tmp mfs:29597 15.4M 87.0K 14.6M 1% /var/run mfs:35201 30.9M 113K 29.3M 1% /var/log ``` Installation went smoothly, except the RPi5 wouldn't boot without a serial adapter plugged into the JST port. RX pin floated with no internal pull-up enabled. Resolved by adding an external 10 kΩ pull-up (RX to 3.3 V). Worked well for two weeks. Then the file system corrupted twice within a week. Disk had 400M+ free space; badblocks revealed no bad sectors. fsck_ffs reported a partially truncated inode indicating an interrupted write; momentary link glitch perhaps—no way to know now. Relocated /dev (sshd pty allocation) to mfs. Moved cron's log to /var/log and symlinked ntpd.drift to /var/run/ntpd.drift. Assigned a static IP and disabled most daemons (dhcpleased, resolvd, slaacd). / is now read-only: ``` /dev/sd0a on / type ffs (local, noatime, wxallowed, read-only) mfs:69255 on /dev type mfs (asynchronous, local, nosuid, size=32768 512-blocks) mfs:75124 on /tmp type mfs (asynchronous, local, nodev, nosuid, size=131072 512-blocks) mfs:15499 on /var/log type mfs (asynchronous, local, nodev, nosuid, size=65536 512-blocks) mfs:14160 on /var/run type mfs (asynchronous, local, nodev, nosuid, size=32768 512-blocks) ``` Security: PGP keys authenticate requests. lex-cgi (chrooted), lex-llm, lex-mailer employ pledge and unveil: ``` my $pid = fork(); if ($pid == 0) { pledge(qw(stdio proc exec)); # sendmail } # unveil rules... unveil($dict_dir, 'r'); unveil(); # unveil lock # main process drops 'exec' pledge(qw(stdio rpath wpath cpath proc)); ``` Packet filter blocks network scans, access to other hosts: ``` set skip on lo block all block out on $net_if to $home_net pass out on $net_if proto { tcp, udp } to $router port domain pass out on $net_if inet proto tcp from ($net_if) to $vps port ssh pass out on $net_if proto udp to ! $home_net port ntp pass out on $net_if proto tcp to ! $home_net port https pass in on $net_if inet proto tcp from $home_net to ($net_if) port ssh ``` End of line. Source: b904341