The Architecture I Built vs. The Architecture I Could Deploy
I participated in a competition and it gave me an insight into something interesting and it led me to build Vicquant. An offline financial advsior app but when i was building it, i ran into a problem and and this aid me in solving real world problem and be innovative for the time being. Vicquant's core idea is a dual local-model setup with zero cloud dependency: Llama-3.1-8B-Instruct for deep…
The competition inspired author to create Vicquant, an offline financial advisor app. However, when building the app, the author encountered a problem in deploying the local models. The solution was to design a dual local-model setup with zero cloud dependency, using Llama-3.1-8B-Instruct for deep financial reasoning and Phi-3.5-mini as a fast, low-RAM fallback. This design was real, with code and model-routing logic existing and passing testing on an 8GB RAM/4 vCPU target at roughly 12-15 tokens/second.
Currently, the version of Vicquant live does not run the local models in production due to the hosting plan not being purchased. Instead, it is connected to OpenRouter, allowing the app to serve AI responses today. This means the deployed demo requires internet and routes model calls through OpenRouter, with offline mode disabled. The local-inference architecture is still present in the codebase, tested, and will return once the hosting plans can be afforded.
The author writes about this constraint to maintain credibility and transparency. Explaining the true situation, even if it results in a less impressive post, is important to avoid discounting other technical claims. Additionally, the gap between the architecture designed and the architecture that can be affordably deployed is a common problem worth discussing.
The author's explanation, whether generated by a local Phi-3.5 model or an OpenRouter-hosted model, remains the same - to describe an already-correct number in plain language, ensuring the backtest gate is passed before any information reaches the user.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.