🐗⚡ Building BOAR | How I Fit an Offline AI Research Engine Into a Phone

1500x500.jpeg

"It all started with a suspicious link sent by a hacker named Vlad."

Building BOAR: How We Fit an Offline AI Research Engine Into a Phone

If you've been following me you know I love testing new things, pushing limits, and bringing dev energy into real world projects.

Whether its building terminal tools, testing Web3 in daily life or coding late into the night, the goal is sovereignty, autonomy, and building tools that works.

image.png

It all started Last Week, @xvlad sent me a DM with just a link...
You know me, I don't click on links, but hey, its from Vlad and poidh.xyz, I know that website...

So, looks like Kenny posted a bounty for building a true offline AI research app for Android: one that can answer complex research questions with zero internet signal, running inside strict limits of 12 GB of RAM and 50 GB of storage.

Mission accepted!!! Lets build!

It sounded as simple as "just put llama.cpp inside a mobile app"... It wasn't. The phone fought back at every single step.

Here is the story of how we built BOAR, the bloopers, the thermal throttling, and what happens when you run local LLMs in airplane mode.


1. From AOAIR to BOAR

The project actually started as AOAIR (Android Offline AI Researcher). But during testing, I kept looking at the name and reading it as "OAIR". Then my brain made the jump: BOAR.

Suddenly I could picture a wild boar out in the middle of nowhere: rugged, curious, independent, moving through the wild without needing anyone's Wi-Fi.

The idea stuck! Before the code was even finished, the mascot already existed in my head.

And that's how AOAIR became BOAR.

2. One Brain at a Time & The 4.5-Second Tax

Rule #1 of mobile hardware: a phone can't keep multiple AI models warm in memory at the same time. BOAR keeps exactly one model loaded. Switching models means unloading, loading from disk, and only then generating.

So instead of routing questions between models, BOAR routes by depth. The model you picked always answers. A confident lookup can be answered instantly from a source sentence with no LLM at all, a normal question goes to your model over a compact set of sources, and a "go deeper" answer uses a larger model only when one is installed and fits in memory.

We learned that the hard way: on the phone, every model switch cost 4.5 to 8.9 seconds. On a phone, which model is already warm in RAM matters just as much as which model is the smartest.

3. Why black holes bend spacetime?

1.gif

One of our very first search tests was simple enough: "Why do black holes bend spacetime?"

Our local RAG engine located the relevant Wikipedia chunks instantly and BOAR delivered a crisp, accurate response. SUCCESS!

...But there is always a "but", isn't there? 🙃

When we pushed the engine with broader, multi-layered scientific queries like "How Do Vaccines Work?", the initial search setup hit a wall:

  • Vaccine ✅ (Found exact word match)
  • Immune system ✅ (Found exact word match)
  • CRISPR & Gene Therapy ⛔
  • Tim Peto / Key Researchers ⛔
  • Complementarity-determining regions ⛔

The early search engine was treating the prompt as one rigid phrase rather than understanding the underlying biological web of concepts. It either matched exact words or left crucial context out entirely.

The Fix:

To solve this, I completely rebuilt our RAG pipeline into a hybrid engine combining ultra-fast keyword relevance (SQLite FTS5) with vector semantic search. We also introduced smart context packing to ensure no single document crowded out the LLM's context window.

However, running this on everyday mobile hardware meant facing a brutal reality:

every vector similarity check comes with a heavy CPU and battery cost.

Balancing deep retrieval precision against battery health and sub-5-second offline response times is a constant tightrope walk when you’re building sovereign AI off-grid.

4. First full benchmark run

Instead of relying on "it feels fast," we built an automated on-device benchmark: 17 research questions tested across models.

On our first full benchmark run on a Dimensity 8300 phone with 11.6 GB of RAM:

  • 85 answers generated in 66 minutes. (Lets improve this number...)
  • The small model (Qwen2.5-1.5B) ran at 11.4 tokens/sec (on a hot phone).
  • Thermal throttling is real: the phone heated up from 39 °C to 43 °C during the run, slowing down later models. Server benchmarks never see that!
  • Reading takes longer than writing: the real bottleneck wasn't token generation speed. It was reading the retrieved context before generating the TTFT (Time To First Token).

5. 50,000 Articles in Your Pocket

2.gif

To give the app deep offline knowledge without melting the phone's CPU, we pre-indexed 49,832 Wikipedia Vital Articles (115,813 passages) on a desktop into a vector index.

The result is a 164 MB file where laptop and phone vectors match at 0.9999 similarity. Combined with this pack, our default 1.5B model scored 6/6 on complex knowledge tests in ~10 seconds each in airplane mode!

Users can also drop their own notes into "My Documents" to build custom local knowledge bases right on their phone.

This comes directly from my own experience on the road.

As a traveler, you’re always collecting pieces of information, itineraries, local guides, emergency contacts, and saved blogs.

But without roaming data or Wi-Fi, those files are just dead weight in a downloads folder.

By allowing users to drop their own notes and PDFs directly into 'My Documents', BOAR turns your phone into a custom local knowledge base.

You can query your own travel research instantly without sending a single byte of data to the cloud. Privacy matter.

6. Android Bloopers from the War Room

Building for physical Android hardware meant fighting the operating system:

  • HyperOS killed model downloads whenever the screen turned off.
  • The "Stop" button took 6 minutes to stop during heavy Llama 3.1 prompt processing.
  • Tone settings ("Summary" vs. "Detailed") were initially ignored because the instruction was buried under 2,000 characters of retrieved text. Moving the style directive next to the user query solved it immediately.
  • 7 engine builds: about 70 MB of our 122 MB APK is llama.cpp compiled into 7 different ARM CPU target versions, so every phone gets its optimal chip instructions (i8mm, dotprod).

7. The Wild 24 Hours 🐗⛺️

After shipping v1.0.0 and submitting it to the bounty, things moved very quickly.

While I went to sleep and rested after 3 days of coding non-stop:

  1. Vitalik Buterin found out the poidh.xyz bounty and he installed and tested BOAR on a physical phone.
  2. People on X discovered the Clanker mascot on Farcaster and started asking whether it was real or connected to the project.
  3. We decided to keep the token out of the main repo because of the bounty specifications, and to keep the core open source project clean.
  4. That didnt stop the story from spreading. The mascot went viral, and the BOAR token narrative took off around the project.
  5. Within roughly 24 hours, a community had formed around something that had started as a small offline AI experiment.

When our designer rendered the BOAR mascot for the first time, I fell in love with it. It was just way too cute!.

Since launching tokens on Clanker is so easy now, I got super excited, sent the mascot art to Vlad: *"Look at this mascot... clank it?"

Vlad took one look at it and said: "Clankit!"

So, we clanked it and minted it right there.

Right after, I went to my dev agent to tell him about the new token, and he immediately hit me with a reality check:

"Keep the repo 100% open source and keep all token mechanics out of the codebase."

He was totally right. I took a step back, respected the boundary, and kept the core GitHub repository strictly focused on open source tech and went to bed.

According to @mengao, maybe that was something that made it catch the hype, that we didn't launch the token as part of the core product. Who knows?

The token remains a fun, community minted artifact on the side, while the BOAR engine itself stays 100% clean and open source, focusing purely on the technology.

8. Seven Times Faster to the First Token

That 11.4 tokens/sec from section 4 was honest, but unfair. The phone was juggling five models and heating up the whole time, and the same model had already done 16–20 tok/s in a cooler test.

As soon as we noticed how much heat was moving the numbers, we started measuring the phone's temperature with our runs.

What heat does to a benchmark:

Run (POCO X6 Pro, 12 GB, Dimensity 8300)Battery temperatureQwen2.5-1.5B
5-model run, 66 minutes, models switching39 °C → 43 °C11.4 tok/s
LFM2.5-8B-A1B run, 17 answers34.5 °C → 42.7 °C(not in this run)
Knowledge-pack test the same evening, model kept loadednot recorded16.8 tok/s

Updates from time reviewing this post:

Same phone, same model, same day: the difference is heat and model switching, not code. So instead of chasing raw speed, we went after what you actually feel: waiting for the first token.

Same phone, same 17 questions, Qwen2.5-1.5B (medians):

DateTokens/secFirst tokenFull answer
24 Sep, first benchmark11.413.6 s21.0 s
28 Sep17.11.9 s5.9 s
29 Sep, first signed shared run18.3–score 95/100

How we got the first token from 13.6 s to 1.9 s:

  • BOAR now sends the model only the passages that actually match the question.
  • Only the single best passage gets filled in with its surrounding sentences.
  • Small talk like "hey, what's up?" skips the search entirely.
  • The start of the answer prompt stays pre computed between chats.
  • The chat stopped redrawing every message on every token, which leaves more CPU for the model.

And it isn't just about lots of RAM: a POCO F3 with 8 GB Snapdragon 870 (a 2021 phone), runs the same model at 19 tok/s, with the first token in 2.6 s and a full answer in 7.2 s.

I honestly wasn’t expecting that... my old phone is still a beast. haha 🐗⚡

3.gif

The full model lineup from our first benchmark (POCO X6 Pro, 12 GB, before the speed work above):

ModelArchitectureTokens/secFirst tokenPeak memoryAnswers completed
LFM2.5-8B-A1BMoE, 8B total, ~1.5B active14.827.2 s5.2 GB17/17
Qwen2.5-1.5Bdense, 1.5B11.413.6 s3.1 GB17/17
Phi-3.5-minidense, 3.8B4.044.0 s4.8 GB13/17
Qwen2.5-7Bdense, 7B2.7~68 s5.1 GB12/17
Instella-MoE-16B-A3BMoE, 16B total–––0/17 (architecture not supported by this llama.cpp build)

The mixture-of-experts model writes as fast as a 1.5B dense model while carrying 8B parameters of knowledge. That's the direction we're pushing: bigger models whose weights mostly sit on disk, with only a small part working on each token.

Run it yourself. The benchmark is in the app (Performance → Evaluation), and from a computer with the phone on USB:

npm run eval:device -- --models qwen2.5-1.5b-instruct-q4km

The best part: these numbers will soon come from the community. Anyone can run the benchmark in the app and share it. Every shared run will be signed by a key inside the phone's own security chip, so the public leaderboard only holds real phones, not scripts. Our first signed run scored 95/100.

9. Experimenting with Disk-Streaming & Massive MoEs 🧪

After the main release, we couldn't resist pushing the hardware even further.

I compiled Colibrì, a pure C engine that streams Mixture-of-Experts weights directly from disk to see if we could run models that exceed normal phone RAM.

What I Learned:

  • Streaming is a game-changer when models exceed RAM: When a model is too big for physical memory, standard llama crawled at 0.07 tok/s. Colibrì's streaming approach was ~25x faster while using less than half the RAM.
  • When models fit in RAM, llama cpp wins easily: On OLMoE that fits in RAM, llama.cpp hit 21.6 tok/s vs. 1.8 tok/s on streaming.
  • A real 35B model actually runs on a phone: It crawled at ~1.5 tok/s, but seeing a 35-billion parameter model generate coherent answers on a mobile device without any cloud connection was mind-blowing.
  • The bottleneck is memory swapping, not storage or heat: Engine diagnostics revealed ~8,400 major page faults per token. Our 5 GB expert cache allocation was simply too large for the phone's free RAM budget, forcing the OS to constantly swap pages back and forth.

Why We Parked It (For Now):

To be completely honest, this initial test was a late-night, caffeinated "let's see if this even compiles on arm64" session.

I was running low on sleep, juggling memory logs, and trying to see how far we could stretch the limits.

We parked it because 1.5 tok/s isn't fast enough for daily research, and for any model that fits in RAM, llama.cpp is still the undisputed speed king.

But this is far from over, Once I get a clear head and a proper block of dev time, I want to dive back into the C codebase, tune the expert cache allocation, optimize mmap, and see if we can tame those page faults!

What we measured on the POCO X6 Pro (Dimensity 8300, 11.6 GB RAM, UFS 4.0):

Engine and modelWhere the weights liveGenerationPeak RAM
llama.cpp, OLMoE-1B-7B Q4_K_Mall in RAM21.6 tok/s~4.2 GB
colibri, OLMoE-1B-7B int8, 16 of 64 experts cachedstreamed from storage~1.8 tok/s3.0 GB
llama.cpp, OLMoE-1B-7B Q8_0bigger than free RAM0.07 tok/sfills RAM
BigMoeOnEdge, Qwen3.6-35B-A3B Q2_K_XL (12.3 GB file)streamed from storage1.25 tok/s (cold cache)~5 GB

If that gets the 35B above 10 tok/s, the end goal is for BOAR to pick the engine on its own: llama.cpp when the model fits in RAM, the streaming engine for a MoE bigger than RAM, with no setting to fiddle with. If you've squeezed big MoE models onto phones, come build it with us.

I want to get back to this and @mengao wants to join...


What's Next?

4.gif

BOAR is 100% open source, works fully offline after a one-time model download, and is built for privacy and resilience.

We're continuing to expand native iOS support, on-device voice STT/TTS, and offline P2P mesh sharing, while constantly optimizing inference and pushing to run larger MoE models on phones with better performance and efficiency.

The goal is simple: keep getting better results from the hardware already in your pocket.

Check out the repo, run the benchmark on your device, and let's keep building parallel, sovereign tools. ⛺⚡

👉 GitHub: rferrari/boar-app


🔗 Reference Links & Resources


1500x500.jpeg

Now the BOAR lore is written forever on HIVE blockchain. 🐗⚡


0
0
0.000
0 comments