1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
|
---
title: 'Lex: LLM dictionary on RPi5'
date: 2026-09-20
layout: post
---
Year of our Lord 2026. LLMs are everywhere.
Built Lex, a context-aware dictionary.
<table style="width: 100%; border-collapse: collapse;">
<tr>
<td>
<img src="rpi5.jpg" alt="RPi5" style="max-height: 450px; width: auto; max-width: 100%;">
</td>
<td>
<img src="lex-email.png" alt="Email" style="max-height: 450px; width: auto; max-width: 100%;">
</td>
</tr>
</table>
Lex comprises three subsystems: lex-cgi, lex-llm, lex-mailer. User emails a
word and an optional context (e.g., define quaint in 'quaint river town').
lex-cgi receives webhook, runs lex-llm, writes definition to file. lex-mailer
emails past entries daily, timed for spaced repetition.
lex-cgi and lex-mailer could run on a $5/mo VPS. LLM inference, however, is
memory intensive. A monolith hosting all three components would entail
$850-$950 billion chatbots, data centers in space, and risking
apocalypse—overkill for defining the occasional word. Even the less glamorous
local inference options require a $20-$40/mo VPS. Decoupling them would let
lex-llm run on hardware I own—a Raspberry Pi 5 I happened to have lying around.
lex-cgi: authenticates PGP signature, sanitizes input, and spools request to
disk:
```
my ($pgp_sig) = $sig_part_raw =~ /(-----BEGIN PGP SIGNATURE-----[\s\S]*?-----END PGP SIGNATURE-----)/;
verify_pgp_signature($signed_part, $pgp_sig)
my $plain_text = $parts[0]->body_str // '';
open($sfh, '>:utf8', $tmp_spool_file);
print $sfh $plain_text . "\n";
close($sfh);
rename($tmp_spool_file, $spool_file);
```
A shell daemon periodically fetches them via SSH, runs lex-llm (listening on a
Unix domain socket), and syncs output back to VPS:
```
prompt=$(ssh -F "${SSH_CONFIG}" -n "${SSH_HOST}" \
"cat '${SPOOL_DIR}/${SPOOL_FILE}'")
response=$(print -r -- "${prompt}" \
| nc -w 300 -U "${SOCK_PATH}" \
| awk '# formatting fix, omitted'
)
# ${remote_cmd} validates output, writes to
# tmp_file, atomically mv into place, rm spool file
print -r -- "${response}" \
| fold -s -w 72 \
| ssh -F "${SSH_CONFIG}" "${SSH_HOST}" "${remote_cmd}"
)
```
lex-mailer: emails one definition daily, chosen for spaced repetition with fair
scheduling:
```
my $pid = fork();
if ($pid == 0) {
# child proc with its own pledge
pledge(qw(stdio proc exec));
my $raw_email;
{
local $/;
$raw_email = <$child_sock>;
}
open(my $mail_pipe, '|-', $sendmail, '-i', '-f', $from, $to);
binmode($mail_pipe, ':utf8');
print $mail_pipe $raw_email;
exit 0;
}
# Candidate selection (spaced repetition + fair scheduling)
my @due_files = grep { $state{$_}{next_due} le $today } @current_files;
my %weights;
my $total_weight = 0;
foreach my $file (@due_files) {
my $w = 1.0 / ($state{$file}{reviews} + 1);
$weights{$file} = $w;
$total_weight += $w;
}
# Weighted random selection (roulette wheel algorithm)
my $rand_point = rand($total_weight);
my $accum = 0;
foreach my $file (@due_files) {
$accum += $weights{$file};
if ($rand_point <= $accum) {
$selected_file = $file;
last;
}
}
```
lex-llm: runs llama.cpp + LLM inference to generate context- and
tone-appropriate definitions:
```
for (i = 0; i < MAX_TOKENS; i++) {
new_token_id = llama_sampler_sample(smpl, ctx, -1);
llama_sampler_accept(smpl, new_token_id);
if (llama_vocab_is_eog(vocab, new_token_id))
break;
char buf[128];
n = llama_token_to_piece(vocab, new_token_id, buf, sizeof(buf), 0, false);
if (n > 0) {
fwrite(buf, 1, (size_t)n, out);
fflush(out);
}
batch = llama_batch_get_one(&new_token_id, 1);
if (llama_decode(ctx, batch) != 0)
break;
}
```
Prompts that took 12 seconds on T490 takes 55 seconds on the RPi5's slower
Cortex-A76. Caching tokens from the previous prompt and decoding only the part
that changes reduced that to 35 seconds:
```
/* drop everything after common prefix. */
llama_memory_seq_rm(llama_get_memory(llm->ctx), 0, n_common, -1);
/* decode the diff */
if (n_common < n_new_tokens) {
batch = llama_batch_get_one(new_tokens + n_common,
n_new_tokens - n_common);
if (llama_decode(llm->ctx, batch) != 0)
return;
memcpy(llm->cached_tokens + n_common,
new_tokens + n_common,
(size_t)(n_new_tokens - n_common) * sizeof(llama_token));
llm->n_cached_tokens = n_new_tokens;
}
```
Forcing the model to RAM at init via LLAMA_LOAD_MODE_MLOCK would have been
desirable, but RPi5 lacks RAM to grant the mlock. LLAMA_LOAD_MODE_NONE yielded
better steady-state performance than LLAMA_LOAD_MODE_MMAP (1.3 GB less MAX RSS,
1-2 seconds faster).
Deploying an LLM to a RPi5 with 4 GB RAM and 3.5 GB flash storage is its own
adventure. Experimented with different models (Q4_K_M) and picked the smallest
viable model first:
<pre class="pre-no-style">
+--------------+--------+-------------+---------------------+----------+
| Model | Tokens | Tok/s (μ±σ) | Time s (Med/Range) | RAM (MB) |
+--------------+--------+-------------+---------------------+----------+
| Qwen 2.5 3B | 114.0 | 5.22 ± 0.31 | 21.79 (19.39-25.20) | 3,235.22 |
| Qwen 2.5 7B | 104.5 | 2.09 ± 0.07 | 49.27 (45.50-52.98) | 7,335.18 |
| Phi 3.5 3.8B | 128.2 | 2.06 ± 0.10 | 71.87 (48.67-72.37) | 3,922.55 |
| LLaMA 3.1 8B | 115.5 | 2.01 ± 0.08 | 57.10 (53.95-66.09) | 7,873.52 |
| Mistral 7B | 106.6 | 1.67 ± 0.13 | 64.25 (57.85-66.86) | 7,535.61 |
+--------------+--------+-------------+---------------------+----------+
</pre>
Larger models caught nuances smaller ones missed, but definitions weren't
consistently better. Gemma 2 9B generated empty responses due to chat template
bug; fixed, but 7.8 GB RAM usage didn't justify re-running benchmarks. Qwen 2.5
3B struggled with formality level, but produced tighter, faster definitions at
3.2 GB RAM and 2 GB storage.
Flashed OpenBSD onto the Toshiba micro SD card without man, game, and X file
sets. Kept comp set to build lex-llm. lex-llm requires just over 2.2 GB to
build. 2.2 GB left on the disk. Disabled boot-time library re-linking and
purged /usr/share/relink to reclaim a further 450 MB.
Stripped llama.cpp to essentials during the build to yield a 12.9 MB
executable:
```
cmake -S $(LLAMA_SRC_DIR) -B $(LLAMA_BUILD) \
-DBUILD_SHARED_LIBS=OFF \
-DGGML_OPENMP=OFF \
-DGGML_VULKAN=OFF \
-DLLAMA_BUILD_SERVER=OFF \
-DLLAMA_BUILD_UI=OFF
# + misc size-reduction flags, omitted
cmake --build $(LLAMA_BUILD) --config Release -j$$(sysctl -n hw.ncpu)
```
Backed up the executable and reinstalled OS removing also the comp set. OS +
lex-llm + model now fits:
```
Filesystem Size Used Avail Capacity Mounted on
/dev/sd0a 3.5G 2.5G 815M 76% /
mfs:76650 61.9M 5.0K 58.8M 1% /tmp
mfs:29597 15.4M 87.0K 14.6M 1% /var/run
mfs:35201 30.9M 113K 29.3M 1% /var/log
```
Worked well for two weeks. Then the file system corrupted twice within a week.
Disk had 400 MB+ free space; badblocks revealed no bad sectors. fsck_ffs
reported a partially truncated inode indicating an interrupted write.
Power-induced link glitch perhaps—no way to know now.
Hunted down anything that wrote to disk (cron, ntpd, sshd). Mounted /dev (sshd
pty allocation) to mfs. Moved cron's log to /var/log/cron and symlinked
ntpd.drift to /var/run/ntpd.drift. Assigned a static IP and disabled almost all
the daemons (dhcpleased, resolvd, slaacd). Placed a /root/.marker to catch
anything I might have missed. / is now read-only:
```
/dev/sd0a on / type ffs (local, noatime, wxallowed, read-only)
mfs:69255 on /dev type mfs (asynchronous, local, nosuid, size=32768 512-blocks)
mfs:75124 on /tmp type mfs (asynchronous, local, nodev, nosuid, size=131072 512-blocks)
mfs:15499 on /var/log type mfs (asynchronous, local, nodev, nosuid, size=65536 512-blocks)
mfs:14160 on /var/run type mfs (asynchronous, local, nodev, nosuid, size=32768 512-blocks)
```
Deployment went smoothly, except the RPi5 wouldn't boot without a serial
adapter plugged into the JST port; RX pin floated with no internal pull-up
enabled. Resolved by adding an external 10 kΩ pull-up (RX to 3.3 V).
Security: PGP authentication, lockfile semaphores, chroot, pledge + unveil
guard the components. Bringing lex-llm into home network has inherent risks.
In addition to the hardened OS, configured packet filter to block network
scans, access to other hosts:
```
set skip on lo
block all
block out on $net_if to $home_net
# DNS lookups for ntpd
pass out on $net_if proto { tcp, udp } to $router port domain
# outbound ssh to VPS
pass out on $net_if inet proto tcp from ($net_if) to $vps port ssh
# ntpd constraints check
pass out on $net_if proto udp to ! $home_net port ntp
pass out on $net_if proto tcp to ! $home_net port https
# inbound SSH from local network management hosts
pass in on $net_if inet proto tcp from $home_net to ($net_if) port ssh
```
End of line.
Source: <a
href="https://git.asciimx.com/lex/tree/?id=b904341947bfeac99fe91d3c9ace760f7da312fd"
class="external" target="_blank" rel="noopener noreferrer">b904341</a>
|