diff options
| author | Sadeep Madurange <sadeep@asciimx.com> | 2026-09-20 12:07:23 +0800 |
|---|---|---|
| committer | Sadeep Madurange <sadeep@asciimx.com> | 2026-09-21 09:54:31 +0800 |
| commit | 612c9d19e683591cfd3b13fcb350e94874157439 (patch) | |
| tree | 72d775ba30870a9a0c5ef2ae0215fd2e50a6f80e /_log | |
| parent | cbe73bd012fbf5aa206c19450e46d76927283392 (diff) | |
| download | www-612c9d19e683591cfd3b13fcb350e94874157439.tar.gz | |
Lex rewrite.minimalist
Diffstat (limited to '_log')
| -rw-r--r-- | _log/lex.md | 306 |
1 files changed, 157 insertions, 149 deletions
diff --git a/_log/lex.md b/_log/lex.md index adebd49..e4c4e0c 100644 --- a/_log/lex.md +++ b/_log/lex.md @@ -1,12 +1,12 @@ --- title: 'Lex: LLM dictionary on RPi5' -date: 2026-09-19 +date: 2026-09-20 layout: post --- Year of our Lord 2026. LLMs are everywhere. -Built Lex, a context-aware dictionary. +Built Lex, a context-aware dictionary. <table style="width: 100%; border-collapse: collapse;"> <tr> @@ -19,29 +19,104 @@ Built Lex, a context-aware dictionary. </tr> </table> -Components: lex-cgi, lex-llm, lex-mailer. User emails word and context. +Lex comprises three subsystems: lex-cgi, lex-llm, lex-mailer. User emails a +word and an optional context (e.g., define quaint in 'quaint river town'). lex-cgi receives webhook, runs lex-llm, writes definition to file. lex-mailer emails past entries daily, timed for spaced repetition. -Tried local inference using llama.cpp + Llama 3.1 8B (GGUF) first: +lex-cgi and lex-mailer could run on a $5/mo VPS. LLM inference, however, is +memory intensive. A monolith hosting all three components would entail +$850-$950 billion chatbots, data centers in space, and risking +apocalypse—overkill for defining the occasional word. Even the less glamorous +local inference options require a $20-$40/mo VPS. Decoupling them would let +lex-llm run on hardware I own—a Raspberry Pi 5 I happened to have lying around. + +lex-cgi: authenticates PGP signature, sanitizes input, and spools request to +disk: + +``` +my ($pgp_sig) = $sig_part_raw =~ /(-----BEGIN PGP SIGNATURE-----[\s\S]*?-----END PGP SIGNATURE-----)/; +verify_pgp_signature($signed_part, $pgp_sig) + +my $plain_text = $parts[0]->body_str // ''; + +open($sfh, '>:utf8', $tmp_spool_file); +print $sfh $plain_text . "\n"; +close($sfh); + +rename($tmp_spool_file, $spool_file); +``` + +A shell daemon periodically fetches them via SSH, runs lex-llm (listening on a +Unix domain socket), and syncs output back to VPS: ``` -llama_backend_init(); -mparams = llama_model_default_params(); -mparams.n_gpu_layers = 0; /* force all layers onto CPU */ -model = llama_model_load_from_file(MODEL_PATH, mparams); - -cparams = llama_context_default_params(); -cparams.n_ctx = N_CTX; /* 1024 */ -llama_init_from_model(model, cparams); +prompt=$(ssh -F "${SSH_CONFIG}" -n "${SSH_HOST}" \ + "cat '${SPOOL_DIR}/${SPOOL_FILE}'") + +response=$(print -r -- "${prompt}" \ + | nc -w 300 -U "${SOCK_PATH}" \ + | awk '# formatting fix, omitted' +) + +# ${remote_cmd} validates output, writes to +# tmp_file, atomically mv into place, rm spool file +print -r -- "${response}" \ + | fold -s -w 72 \ + | ssh -F "${SSH_CONFIG}" "${SSH_HOST}" "${remote_cmd}" +) ``` -Generation loop synthesizes definition: +lex-mailer: emails one definition daily, chosen for spaced repetition with fair +scheduling: + +``` +my $pid = fork(); +if ($pid == 0) { + # child proc with its own pledge + pledge(qw(stdio proc exec)); + + my $raw_email; + { + local $/; + $raw_email = <$child_sock>; + } + open(my $mail_pipe, '|-', $sendmail, '-i', '-f', $from, $to); + binmode($mail_pipe, ':utf8'); + print $mail_pipe $raw_email; + exit 0; +} + +# Candidate selection (spaced repetition + fair scheduling) +my @due_files = grep { $state{$_}{next_due} le $today } @current_files; +my %weights; +my $total_weight = 0; +foreach my $file (@due_files) { + my $w = 1.0 / ($state{$file}{reviews} + 1); + $weights{$file} = $w; + $total_weight += $w; +} + +# Weighted random selection (roulette wheel algorithm) +my $rand_point = rand($total_weight); +my $accum = 0; +foreach my $file (@due_files) { + $accum += $weights{$file}; + if ($rand_point <= $accum) { + $selected_file = $file; + last; + } +} +``` + +lex-llm: runs llama.cpp + LLM inference to generate context- and +tone-appropriate definitions: ``` for (i = 0; i < MAX_TOKENS; i++) { new_token_id = llama_sampler_sample(smpl, ctx, -1); llama_sampler_accept(smpl, new_token_id); + if (llama_vocab_is_eog(vocab, new_token_id)) break; @@ -56,86 +131,11 @@ for (i = 0; i < MAX_TOKENS; i++) { if (llama_decode(ctx, batch) != 0) break; } -``` - -Build what's needed: +``` -``` -cmake -S $(LLAMA_SRC_DIR) -B $(LLAMA_BUILD) \ - -DBUILD_SHARED_LIBS=OFF \ - -DGGML_OPENMP=OFF \ - -DGGML_VULKAN=OFF \ - -DLLAMA_BUILD_SERVER=OFF \ - -DLLAMA_BUILD_UI=OFF - # + misc size-reduction flags, omitted - -cmake --build $(LLAMA_BUILD) --config Release -j$$(sysctl -n hw.ncpu) -``` - -Functional, but requires a $40/mo VPS (8 GB RAM). Benchmarked alternative -Q4_K_M models on T490 (i7-10510U, OpenBSD 7.9, CPU only): - -<pre class="pre-no-style"> -+--------------+--------+-------------+---------------------+----------+ -| Model | Tokens | Tok/s (μ±σ) | Time s (Med/Range) | RAM (MB) | -+--------------+--------+-------------+---------------------+----------+ -| Qwen 2.5 3B | 114.0 | 5.22 ± 0.31 | 21.79 (19.39-25.20) | 3,235.22 | -| Qwen 2.5 7B | 104.5 | 2.09 ± 0.07 | 49.27 (45.50-52.98) | 7,335.18 | -| Phi 3.5 3.8B | 128.2 | 2.06 ± 0.10 | 71.87 (48.67-72.37) | 3,922.55 | -| LLaMA 3.1 8B | 115.5 | 2.01 ± 0.08 | 57.10 (53.95-66.09) | 7,873.52 | -| Mistral 7B | 106.6 | 1.67 ± 0.13 | 64.25 (57.85-66.86) | 7,535.61 | -+--------------+--------+-------------+---------------------+----------+ -</pre> - -Larger models caught nuances smaller ones missed, but definitions weren't -consistently better. Gemma 2 9B generated empty responses due to chat template -bug; fixed, but 7.8 GB RAM usage didn't justify re-running benchmarks. Smaller -models handled shorter prompts better. Qwen 2.5 struggled with formality level, -but produced tighter, faster definitions at 3.2 GB RAM. - -Qwen 2.5 3B halves the VPS cost—still wasteful for the occasional prompt. -Decoupling lex-llm from lex-cgi lets lex-llm run on a 4 GB Raspberry Pi 5; rest -could run on a $5/mo 1 GB VPS. - -Refactored lex-llm to communicate over Unix domain sockets. lex-cgi now spools -incoming emails to disk. A shell daemon periodically fetches them via SSH, runs -lex-llm, and syncs output back to VPS. - -``` -while true; do - # fork to avoid resource leakage on unexpected errors - ( - prompt=$(ssh -F "${SSH_CONFIG}" -n "${SSH_HOST}" \ - "cat '${SPOOL_DIR}/${SPOOL_FILE}'") - - response=$(print -r -- "${prompt}" \ - | nc -w 300 -U "${SOCK_PATH}" \ - | awk '# formatting fix, omitted...' - ) - - # validate output, write to tmp_file, - # atomically mv into place, rm spool file—omitted - - print -r -- "${response}" \ - | fold -s -w 72 \ - | ssh -F "${SSH_CONFIG}" "${SSH_HOST}" "${remote_cmd}" - ) -done -``` - -Performance: lex-llm was 8x slower (190 s/req) on the RPi5. Benchmarked -version loaded the model every request. Loading it once at init reduced -execution time to 55 seconds. - -Forcing the model to RAM at init via LLAMA_LOAD_MODE_MLOCK would have been -desirable, but RPi5 lacks RAM to grant the mlock. Between LLAMA_LOAD_MODE_NONE -and LLAMA_LOAD_MODE_MMAP, former yielded better steady-state performance (1.3 -GB less MAX RSS, 1-2 seconds faster). - -RPi5's slower Cortex-A76 processor was still 2x behind benchmarks. -llama_decode() was evaluating system + user prompt every request. System prompt -doesn't change between requests; cached token offsets and decoded only the part -that changed: +Prompts that took 12 seconds on T490 takes 55 seconds on the RPi5's slower +Cortex-A76. Caching tokens from the previous prompt and decoding only the part +that changes reduced that to 35 seconds: ``` /* drop everything after common prefix. */ @@ -156,63 +156,74 @@ if (n_common < n_new_tokens) { } ``` -Prefix-caching reduced execution time to 35 seconds—acceptable for an -email-driven async dictionary. +Forcing the model to RAM at init via LLAMA_LOAD_MODE_MLOCK would have been +desirable, but RPi5 lacks RAM to grant the mlock. LLAMA_LOAD_MODE_NONE yielded +better steady-state performance than LLAMA_LOAD_MODE_MMAP (1.3 GB less MAX RSS, +1-2 seconds faster). -Note on measurements: 3.4 A supply likely caused CPU throttling on the RPi5 -during measurements. Final version uses the recommended 5 A supply. +Deploying an LLM to a RPi5 with 4 GB RAM and 3.5 GB flash storage is its own +adventure. Experimented with different models (Q4_K_M) and picked the smallest +viable model first: -Fidelity: Didn't notice qualitative differences between T490 and RPi5 output -under normal operation. They exist, however. System prompt fragment duplicated -the SYNONYMS instruction inside GENERAL TONE: +<pre class="pre-no-style"> ++--------------+--------+-------------+---------------------+----------+ +| Model | Tokens | Tok/s (μ±σ) | Time s (Med/Range) | RAM (MB) | ++--------------+--------+-------------+---------------------+----------+ +| Qwen 2.5 3B | 114.0 | 5.22 ± 0.31 | 21.79 (19.39-25.20) | 3,235.22 | +| Qwen 2.5 7B | 104.5 | 2.09 ± 0.07 | 49.27 (45.50-52.98) | 7,335.18 | +| Phi 3.5 3.8B | 128.2 | 2.06 ± 0.10 | 71.87 (48.67-72.37) | 3,922.55 | +| LLaMA 3.1 8B | 115.5 | 2.01 ± 0.08 | 57.10 (53.95-66.09) | 7,873.52 | +| Mistral 7B | 106.6 | 1.67 ± 0.13 | 64.25 (57.85-66.86) | 7,535.61 | ++--------------+--------+-------------+---------------------+----------+ +</pre> -``` -One terse sentence of describing general tone. One terse sentence of -standard conversational synonyms. -``` +Larger models caught nuances smaller ones missed, but definitions weren't +consistently better. Gemma 2 9B generated empty responses due to chat template +bug; fixed, but 7.8 GB RAM usage didn't justify re-running benchmarks. Qwen 2.5 +3B struggled with formality level, but produced tighter, faster definitions at +3.2 GB RAM and 2 GB storage. -RPi5 echoed the second stray sentence verbatim; T490 did not: +Flashed OpenBSD onto the Toshiba micro SD card without man, game, and X file +sets. Kept comp set to build lex-llm. lex-llm requires just over 2.2 GB to +build. 2.2 GB left on the disk. Disabled boot-time library re-linking and +purged /usr/share/relink to reclaim a further 450 MB. -``` -GENERAL TONE: -The phrase suggests the impending end of life with somber and formal -language. -One terse sentence of standard conversational synonyms: Buried or -surrounded by darkness. -``` +Stripped llama.cpp to essentials during the build to yield a 12.9 MB +executable: -Possible artifact of non-associative floating-point arithmetic (x86 SIMD vs ARM -NEON instruction ordering). Dropping duplicate sentence from prompt reconciled -the outputs. +``` +cmake -S $(LLAMA_SRC_DIR) -B $(LLAMA_BUILD) \ + -DBUILD_SHARED_LIBS=OFF \ + -DGGML_OPENMP=OFF \ + -DGGML_VULKAN=OFF \ + -DLLAMA_BUILD_SERVER=OFF \ + -DLLAMA_BUILD_UI=OFF + # + misc size-reduction flags, omitted -Storage: lex-llm binary is 12.9 MB. GGUF takes up 2 GB. Standard OpenBSD -install recommends 8 GB. SD card I have is 4 GB—3.5 GB usable. +cmake --build $(LLAMA_BUILD) --config Release -j$$(sysctl -n hw.ncpu) +``` -Flashed OpenBSD onto SD card; tossed man, game, and X file sets. Kept comp set -to build lex-llm. Disabled boot-time library re-linking and purged -/usr/share/relink to reclaim a further 450 MB. OS + lex-llm + model now fits: +Backed up the executable and reinstalled OS removing also the comp set. OS + +lex-llm + model now fits: ``` Filesystem Size Used Avail Capacity Mounted on -/dev/sd0a 3.5G 2.8G 477M 86% / +/dev/sd0a 3.5G 2.5G 815M 76% / mfs:76650 61.9M 5.0K 58.8M 1% /tmp mfs:29597 15.4M 87.0K 14.6M 1% /var/run mfs:35201 30.9M 113K 29.3M 1% /var/log ``` -Installation went smoothly, except the RPi5 wouldn't boot without a serial -adapter plugged into the JST port. RX pin floated with no internal pull-up -enabled. Resolved by adding an external 10 kΩ pull-up (RX to 3.3 V). - - Worked well for two weeks. Then the file system corrupted twice within a week. -Disk had 400M+ free space; badblocks revealed no bad sectors. fsck_ffs reported -a partially truncated inode indicating an interrupted write; momentary link -glitch perhaps—no way to know now. +Disk had 400 MB+ free space; badblocks revealed no bad sectors. fsck_ffs +reported a partially truncated inode indicating an interrupted write. +Power-induced link glitch perhaps—no way to know now. -Relocated /dev (sshd pty allocation) to mfs. Moved cron's log to /var/log and -symlinked ntpd.drift to /var/run/ntpd.drift. Assigned a static IP and disabled -most daemons (dhcpleased, resolvd, slaacd). / is now read-only: +Hunted down anything that wrote to disk (cron, ntpd, sshd). Mounted /dev (sshd +pty allocation) to mfs. Moved cron's log to /var/log/cron and symlinked +ntpd.drift to /var/run/ntpd.drift. Assigned a static IP and disabled almost all +the daemons (dhcpleased, resolvd, slaacd). Placed a /root/.marker to catch +anything I might have missed. / is now read-only: ``` /dev/sd0a on / type ffs (local, noatime, wxallowed, read-only) @@ -222,35 +233,32 @@ mfs:15499 on /var/log type mfs (asynchronous, local, nodev, nosuid, size=65536 5 mfs:14160 on /var/run type mfs (asynchronous, local, nodev, nosuid, size=32768 512-blocks) ``` -Security: PGP keys authenticate requests. lex-cgi (chrooted), lex-llm, -lex-mailer employ pledge and unveil: - -``` -my $pid = fork(); -if ($pid == 0) { - pledge(qw(stdio proc exec)); - # sendmail -} - -# unveil rules... -unveil($dict_dir, 'r'); -unveil(); # unveil lock - -# main process drops 'exec' -pledge(qw(stdio rpath wpath cpath proc)); -``` +Deployment went smoothly, except the RPi5 wouldn't boot without a serial +adapter plugged into the JST port; RX pin floated with no internal pull-up +enabled. Resolved by adding an external 10 kΩ pull-up (RX to 3.3 V). -Packet filter blocks network scans, access to other hosts: +Security: PGP authentication, lockfile semaphores, chroot, pledge + unveil +guard the components. Bringing lex-llm into home network has inherent risks. +In addition to the hardened OS, configured packet filter to block network +scans, access to other hosts: ``` set skip on lo block all block out on $net_if to $home_net + +# DNS lookups for ntpd pass out on $net_if proto { tcp, udp } to $router port domain + +# outbound ssh to VPS pass out on $net_if inet proto tcp from ($net_if) to $vps port ssh + +# ntpd constraints check pass out on $net_if proto udp to ! $home_net port ntp pass out on $net_if proto tcp to ! $home_net port https + +# inbound SSH from local network management hosts pass in on $net_if inet proto tcp from $home_net to ($net_if) port ssh ``` |
