summaryrefslogtreecommitdiffstats
diff options
context:
space:
mode:
-rw-r--r--_log/lex.md306
1 files changed, 157 insertions, 149 deletions
diff --git a/_log/lex.md b/_log/lex.md
index adebd49..e4c4e0c 100644
--- a/_log/lex.md
+++ b/_log/lex.md
@@ -1,12 +1,12 @@
---
title: 'Lex: LLM dictionary on RPi5'
-date: 2026-09-19
+date: 2026-09-20
layout: post
---
Year of our Lord 2026. LLMs are everywhere.
-Built Lex, a context-aware dictionary.
+Built Lex, a context-aware dictionary.
<table style="width: 100%; border-collapse: collapse;">
<tr>
@@ -19,29 +19,104 @@ Built Lex, a context-aware dictionary.
</tr>
</table>
-Components: lex-cgi, lex-llm, lex-mailer. User emails word and context.
+Lex comprises three subsystems: lex-cgi, lex-llm, lex-mailer. User emails a
+word and an optional context (e.g., define quaint in 'quaint river town').
lex-cgi receives webhook, runs lex-llm, writes definition to file. lex-mailer
emails past entries daily, timed for spaced repetition.
-Tried local inference using llama.cpp + Llama 3.1 8B (GGUF) first:
+lex-cgi and lex-mailer could run on a $5/mo VPS. LLM inference, however, is
+memory intensive. A monolith hosting all three components would entail
+$850-$950 billion chatbots, data centers in space, and risking
+apocalypse—overkill for defining the occasional word. Even the less glamorous
+local inference options require a $20-$40/mo VPS. Decoupling them would let
+lex-llm run on hardware I own—a Raspberry Pi 5 I happened to have lying around.
+
+lex-cgi: authenticates PGP signature, sanitizes input, and spools request to
+disk:
+
+```
+my ($pgp_sig) = $sig_part_raw =~ /(-----BEGIN PGP SIGNATURE-----[\s\S]*?-----END PGP SIGNATURE-----)/;
+verify_pgp_signature($signed_part, $pgp_sig)
+
+my $plain_text = $parts[0]->body_str // '';
+
+open($sfh, '>:utf8', $tmp_spool_file);
+print $sfh $plain_text . "\n";
+close($sfh);
+
+rename($tmp_spool_file, $spool_file);
+```
+
+A shell daemon periodically fetches them via SSH, runs lex-llm (listening on a
+Unix domain socket), and syncs output back to VPS:
```
-llama_backend_init();
-mparams = llama_model_default_params();
-mparams.n_gpu_layers = 0; /* force all layers onto CPU */
-model = llama_model_load_from_file(MODEL_PATH, mparams);
-
-cparams = llama_context_default_params();
-cparams.n_ctx = N_CTX; /* 1024 */
-llama_init_from_model(model, cparams);
+prompt=$(ssh -F "${SSH_CONFIG}" -n "${SSH_HOST}" \
+ "cat '${SPOOL_DIR}/${SPOOL_FILE}'")
+
+response=$(print -r -- "${prompt}" \
+ | nc -w 300 -U "${SOCK_PATH}" \
+ | awk '# formatting fix, omitted'
+)
+
+# ${remote_cmd} validates output, writes to
+# tmp_file, atomically mv into place, rm spool file
+print -r -- "${response}" \
+ | fold -s -w 72 \
+ | ssh -F "${SSH_CONFIG}" "${SSH_HOST}" "${remote_cmd}"
+)
```
-Generation loop synthesizes definition:
+lex-mailer: emails one definition daily, chosen for spaced repetition with fair
+scheduling:
+
+```
+my $pid = fork();
+if ($pid == 0) {
+ # child proc with its own pledge
+ pledge(qw(stdio proc exec));
+
+ my $raw_email;
+ {
+ local $/;
+ $raw_email = <$child_sock>;
+ }
+ open(my $mail_pipe, '|-', $sendmail, '-i', '-f', $from, $to);
+ binmode($mail_pipe, ':utf8');
+ print $mail_pipe $raw_email;
+ exit 0;
+}
+
+# Candidate selection (spaced repetition + fair scheduling)
+my @due_files = grep { $state{$_}{next_due} le $today } @current_files;
+my %weights;
+my $total_weight = 0;
+foreach my $file (@due_files) {
+ my $w = 1.0 / ($state{$file}{reviews} + 1);
+ $weights{$file} = $w;
+ $total_weight += $w;
+}
+
+# Weighted random selection (roulette wheel algorithm)
+my $rand_point = rand($total_weight);
+my $accum = 0;
+foreach my $file (@due_files) {
+ $accum += $weights{$file};
+ if ($rand_point <= $accum) {
+ $selected_file = $file;
+ last;
+ }
+}
+```
+
+lex-llm: runs llama.cpp + LLM inference to generate context- and
+tone-appropriate definitions:
```
for (i = 0; i < MAX_TOKENS; i++) {
new_token_id = llama_sampler_sample(smpl, ctx, -1);
llama_sampler_accept(smpl, new_token_id);
+
if (llama_vocab_is_eog(vocab, new_token_id))
break;
@@ -56,86 +131,11 @@ for (i = 0; i < MAX_TOKENS; i++) {
if (llama_decode(ctx, batch) != 0)
break;
}
-```
-
-Build what's needed:
+```
-```
-cmake -S $(LLAMA_SRC_DIR) -B $(LLAMA_BUILD) \
- -DBUILD_SHARED_LIBS=OFF \
- -DGGML_OPENMP=OFF \
- -DGGML_VULKAN=OFF \
- -DLLAMA_BUILD_SERVER=OFF \
- -DLLAMA_BUILD_UI=OFF
- # + misc size-reduction flags, omitted
-
-cmake --build $(LLAMA_BUILD) --config Release -j$$(sysctl -n hw.ncpu)
-```
-
-Functional, but requires a $40/mo VPS (8 GB RAM). Benchmarked alternative
-Q4_K_M models on T490 (i7-10510U, OpenBSD 7.9, CPU only):
-
-<pre class="pre-no-style">
-+--------------+--------+-------------+---------------------+----------+
-| Model | Tokens | Tok/s (μ±σ) | Time s (Med/Range) | RAM (MB) |
-+--------------+--------+-------------+---------------------+----------+
-| Qwen 2.5 3B | 114.0 | 5.22 ± 0.31 | 21.79 (19.39-25.20) | 3,235.22 |
-| Qwen 2.5 7B | 104.5 | 2.09 ± 0.07 | 49.27 (45.50-52.98) | 7,335.18 |
-| Phi 3.5 3.8B | 128.2 | 2.06 ± 0.10 | 71.87 (48.67-72.37) | 3,922.55 |
-| LLaMA 3.1 8B | 115.5 | 2.01 ± 0.08 | 57.10 (53.95-66.09) | 7,873.52 |
-| Mistral 7B | 106.6 | 1.67 ± 0.13 | 64.25 (57.85-66.86) | 7,535.61 |
-+--------------+--------+-------------+---------------------+----------+
-</pre>
-
-Larger models caught nuances smaller ones missed, but definitions weren't
-consistently better. Gemma 2 9B generated empty responses due to chat template
-bug; fixed, but 7.8 GB RAM usage didn't justify re-running benchmarks. Smaller
-models handled shorter prompts better. Qwen 2.5 struggled with formality level,
-but produced tighter, faster definitions at 3.2 GB RAM.
-
-Qwen 2.5 3B halves the VPS cost—still wasteful for the occasional prompt.
-Decoupling lex-llm from lex-cgi lets lex-llm run on a 4 GB Raspberry Pi 5; rest
-could run on a $5/mo 1 GB VPS.
-
-Refactored lex-llm to communicate over Unix domain sockets. lex-cgi now spools
-incoming emails to disk. A shell daemon periodically fetches them via SSH, runs
-lex-llm, and syncs output back to VPS.
-
-```
-while true; do
- # fork to avoid resource leakage on unexpected errors
- (
- prompt=$(ssh -F "${SSH_CONFIG}" -n "${SSH_HOST}" \
- "cat '${SPOOL_DIR}/${SPOOL_FILE}'")
-
- response=$(print -r -- "${prompt}" \
- | nc -w 300 -U "${SOCK_PATH}" \
- | awk '# formatting fix, omitted...'
- )
-
- # validate output, write to tmp_file,
- # atomically mv into place, rm spool file—omitted
-
- print -r -- "${response}" \
- | fold -s -w 72 \
- | ssh -F "${SSH_CONFIG}" "${SSH_HOST}" "${remote_cmd}"
- )
-done
-```
-
-Performance: lex-llm was 8x slower (190 s/req) on the RPi5. Benchmarked
-version loaded the model every request. Loading it once at init reduced
-execution time to 55 seconds.
-
-Forcing the model to RAM at init via LLAMA_LOAD_MODE_MLOCK would have been
-desirable, but RPi5 lacks RAM to grant the mlock. Between LLAMA_LOAD_MODE_NONE
-and LLAMA_LOAD_MODE_MMAP, former yielded better steady-state performance (1.3
-GB less MAX RSS, 1-2 seconds faster).
-
-RPi5's slower Cortex-A76 processor was still 2x behind benchmarks.
-llama_decode() was evaluating system + user prompt every request. System prompt
-doesn't change between requests; cached token offsets and decoded only the part
-that changed:
+Prompts that took 12 seconds on T490 takes 55 seconds on the RPi5's slower
+Cortex-A76. Caching tokens from the previous prompt and decoding only the part
+that changes reduced that to 35 seconds:
```
/* drop everything after common prefix. */
@@ -156,63 +156,74 @@ if (n_common < n_new_tokens) {
}
```
-Prefix-caching reduced execution time to 35 seconds—acceptable for an
-email-driven async dictionary.
+Forcing the model to RAM at init via LLAMA_LOAD_MODE_MLOCK would have been
+desirable, but RPi5 lacks RAM to grant the mlock. LLAMA_LOAD_MODE_NONE yielded
+better steady-state performance than LLAMA_LOAD_MODE_MMAP (1.3 GB less MAX RSS,
+1-2 seconds faster).
-Note on measurements: 3.4 A supply likely caused CPU throttling on the RPi5
-during measurements. Final version uses the recommended 5 A supply.
+Deploying an LLM to a RPi5 with 4 GB RAM and 3.5 GB flash storage is its own
+adventure. Experimented with different models (Q4_K_M) and picked the smallest
+viable model first:
-Fidelity: Didn't notice qualitative differences between T490 and RPi5 output
-under normal operation. They exist, however. System prompt fragment duplicated
-the SYNONYMS instruction inside GENERAL TONE:
+<pre class="pre-no-style">
++--------------+--------+-------------+---------------------+----------+
+| Model | Tokens | Tok/s (μ±σ) | Time s (Med/Range) | RAM (MB) |
++--------------+--------+-------------+---------------------+----------+
+| Qwen 2.5 3B | 114.0 | 5.22 ± 0.31 | 21.79 (19.39-25.20) | 3,235.22 |
+| Qwen 2.5 7B | 104.5 | 2.09 ± 0.07 | 49.27 (45.50-52.98) | 7,335.18 |
+| Phi 3.5 3.8B | 128.2 | 2.06 ± 0.10 | 71.87 (48.67-72.37) | 3,922.55 |
+| LLaMA 3.1 8B | 115.5 | 2.01 ± 0.08 | 57.10 (53.95-66.09) | 7,873.52 |
+| Mistral 7B | 106.6 | 1.67 ± 0.13 | 64.25 (57.85-66.86) | 7,535.61 |
++--------------+--------+-------------+---------------------+----------+
+</pre>
-```
-One terse sentence of describing general tone. One terse sentence of
-standard conversational synonyms.
-```
+Larger models caught nuances smaller ones missed, but definitions weren't
+consistently better. Gemma 2 9B generated empty responses due to chat template
+bug; fixed, but 7.8 GB RAM usage didn't justify re-running benchmarks. Qwen 2.5
+3B struggled with formality level, but produced tighter, faster definitions at
+3.2 GB RAM and 2 GB storage.
-RPi5 echoed the second stray sentence verbatim; T490 did not:
+Flashed OpenBSD onto the Toshiba micro SD card without man, game, and X file
+sets. Kept comp set to build lex-llm. lex-llm requires just over 2.2 GB to
+build. 2.2 GB left on the disk. Disabled boot-time library re-linking and
+purged /usr/share/relink to reclaim a further 450 MB.
-```
-GENERAL TONE:
-The phrase suggests the impending end of life with somber and formal
-language.
-One terse sentence of standard conversational synonyms: Buried or
-surrounded by darkness.
-```
+Stripped llama.cpp to essentials during the build to yield a 12.9 MB
+executable:
-Possible artifact of non-associative floating-point arithmetic (x86 SIMD vs ARM
-NEON instruction ordering). Dropping duplicate sentence from prompt reconciled
-the outputs.
+```
+cmake -S $(LLAMA_SRC_DIR) -B $(LLAMA_BUILD) \
+ -DBUILD_SHARED_LIBS=OFF \
+ -DGGML_OPENMP=OFF \
+ -DGGML_VULKAN=OFF \
+ -DLLAMA_BUILD_SERVER=OFF \
+ -DLLAMA_BUILD_UI=OFF
+ # + misc size-reduction flags, omitted
-Storage: lex-llm binary is 12.9 MB. GGUF takes up 2 GB. Standard OpenBSD
-install recommends 8 GB. SD card I have is 4 GB—3.5 GB usable.
+cmake --build $(LLAMA_BUILD) --config Release -j$$(sysctl -n hw.ncpu)
+```
-Flashed OpenBSD onto SD card; tossed man, game, and X file sets. Kept comp set
-to build lex-llm. Disabled boot-time library re-linking and purged
-/usr/share/relink to reclaim a further 450 MB. OS + lex-llm + model now fits:
+Backed up the executable and reinstalled OS removing also the comp set. OS +
+lex-llm + model now fits:
```
Filesystem Size Used Avail Capacity Mounted on
-/dev/sd0a 3.5G 2.8G 477M 86% /
+/dev/sd0a 3.5G 2.5G 815M 76% /
mfs:76650 61.9M 5.0K 58.8M 1% /tmp
mfs:29597 15.4M 87.0K 14.6M 1% /var/run
mfs:35201 30.9M 113K 29.3M 1% /var/log
```
-Installation went smoothly, except the RPi5 wouldn't boot without a serial
-adapter plugged into the JST port. RX pin floated with no internal pull-up
-enabled. Resolved by adding an external 10 kΩ pull-up (RX to 3.3 V).
-
-
Worked well for two weeks. Then the file system corrupted twice within a week.
-Disk had 400M+ free space; badblocks revealed no bad sectors. fsck_ffs reported
-a partially truncated inode indicating an interrupted write; momentary link
-glitch perhaps—no way to know now.
+Disk had 400 MB+ free space; badblocks revealed no bad sectors. fsck_ffs
+reported a partially truncated inode indicating an interrupted write.
+Power-induced link glitch perhaps—no way to know now.
-Relocated /dev (sshd pty allocation) to mfs. Moved cron's log to /var/log and
-symlinked ntpd.drift to /var/run/ntpd.drift. Assigned a static IP and disabled
-most daemons (dhcpleased, resolvd, slaacd). / is now read-only:
+Hunted down anything that wrote to disk (cron, ntpd, sshd). Mounted /dev (sshd
+pty allocation) to mfs. Moved cron's log to /var/log/cron and symlinked
+ntpd.drift to /var/run/ntpd.drift. Assigned a static IP and disabled almost all
+the daemons (dhcpleased, resolvd, slaacd). Placed a /root/.marker to catch
+anything I might have missed. / is now read-only:
```
/dev/sd0a on / type ffs (local, noatime, wxallowed, read-only)
@@ -222,35 +233,32 @@ mfs:15499 on /var/log type mfs (asynchronous, local, nodev, nosuid, size=65536 5
mfs:14160 on /var/run type mfs (asynchronous, local, nodev, nosuid, size=32768 512-blocks)
```
-Security: PGP keys authenticate requests. lex-cgi (chrooted), lex-llm,
-lex-mailer employ pledge and unveil:
-
-```
-my $pid = fork();
-if ($pid == 0) {
- pledge(qw(stdio proc exec));
- # sendmail
-}
-
-# unveil rules...
-unveil($dict_dir, 'r');
-unveil(); # unveil lock
-
-# main process drops 'exec'
-pledge(qw(stdio rpath wpath cpath proc));
-```
+Deployment went smoothly, except the RPi5 wouldn't boot without a serial
+adapter plugged into the JST port; RX pin floated with no internal pull-up
+enabled. Resolved by adding an external 10 kΩ pull-up (RX to 3.3 V).
-Packet filter blocks network scans, access to other hosts:
+Security: PGP authentication, lockfile semaphores, chroot, pledge + unveil
+guard the components. Bringing lex-llm into home network has inherent risks.
+In addition to the hardened OS, configured packet filter to block network
+scans, access to other hosts:
```
set skip on lo
block all
block out on $net_if to $home_net
+
+# DNS lookups for ntpd
pass out on $net_if proto { tcp, udp } to $router port domain
+
+# outbound ssh to VPS
pass out on $net_if inet proto tcp from ($net_if) to $vps port ssh
+
+# ntpd constraints check
pass out on $net_if proto udp to ! $home_net port ntp
pass out on $net_if proto tcp to ! $home_net port https
+
+# inbound SSH from local network management hosts
pass in on $net_if inet proto tcp from $home_net to ($net_if) port ssh
```