Scan, Extract & Call
Stop typing numbers manually. Point your camera at business cards, docs, or screens to extract and dial numbers instantly.
Get Scan2Call π±STAKSOFT AI RESEARCH | Scan2PDF Production Engine V1.4
Engineering Whitepaper Β· On-Device Neural Vision
A deep architectural breakdown of zero-cloud mobile document digitization: How we engineered 12β25ms real-time corner regression, non-linear spine unwarping, low-frequency shadow cancellation, and margin inpainting on legacy Android architectures.
Production App
Scan2PDF (v2026.9+)
Aggregate Footprint
~14 MB (4 Quantized Models)
Inference Runtime
Android NNAPI / TFLite GPU
Custom AI Development
Available for Enterprise
Table of Contents
Most consumer scanner apps transmit uncompressed document bitmaps over cellular interfaces to remote clusters running general-purpose segmentation APIs. Staksoftβs Document AI executes the entire vision pipeline directly on local mobile hardware using four dedicated, specialized neural nets (totalling ~14MB in aggregate). This achieves deterministic 12β25ms detection latency, zero byte leakage, complete offline resilience, and zero cloud server operating expense passed on to users.
Engineering Dimension | Cloud-Centric Scanner Pipelines | Scan2PDF On-Device Pipeline |
|---|---|---|
Time-to-First-Crop (TTFC) | 1,800 ms β 6,500 ms (RND + Transfer) | 12 ms β 25 ms (Instantaneous) |
Bandwidth Consumption | 4.5 MB β 18.0 MB per scanned page | 0.00 KB (Zero Transmission) |
Privacy & Attestation | Vulnerable to intermediate MITM, TLS inspection, remote disk logging | Zero External Footprint; strictly compliant with air-gapped security policies |
Infrastructure Scaling Cost | $0.002 to $0.015 per page processed | $0.00 (Computed on local edge silicons) |
Executing computer vision tasks on consumer mobile devices demands zero heap-allocation overhead during the camera preview cycle. Running continuous garbage-collected (GC) memory spikes causes severe UI stutter (jank) in Android's Choreographer frame-rate pipeline.
Camera2 API (YUV_420_888 Frame Buffer) β βββΊ Downsampled Ring-Buffer (256x256x1 Grayscale) βββΊ Model 1: Quad Boundary Net [12-25ms] β β β Calculates [p0, p1, p2, p3] Coordinates β βΌ β Capture Button Triggered / Auto-Stabilization Passed βββββ βΌ Full-Resolution Sensor RAW/JPEG (e.g. 4000x3000) β βββΊ Stage 1: Homography Perspective Transform via Bilinear Interpolation β βββΊ Stage 2: Curvature Classifier Check (Fast 1.2MB Binary Gate) β βββΊ Flat Page Detected βββββββββΊ Bypass Dewarping (0ms overhead) β βββΊ Book Spine Curve Detected βββΊ Model 2: 3D Surface Mesh Net [0.35-0.90s] β βββΊ Stage 3: Auto Color / Lighting Division (Model 3: Residual UNet [0.45s]) β βββΊ Illuminant Field Extraction & Divisive Normalization β βββΊ Stage 4: Margin Inpainting (Model 4: In-Margin Finger Net [0.20s]) β βββΊ Binary Mask Generation βββΊ Contextual Border Smoothing β βΌ Clean Flattened RGB Document βββΊ Offline OCR Pipeline / Structured JSON Engine / Vector PDF
Backbone: Depthwise separable inverted residual blocks derived from MobileNetV2 with width multiplier $\alpha = 0.35$.
Quantized Parameter Size: 2,988,142 bytes (~2.98 MB).
Loss Function: Combined Smooth $L_1$ Regression + Corner Heatmap Dice Loss:
$$\mathcal{L}_{total} = \lambda_1 \sum_{k=1}^{4} \text{Smooth}_{L1}(\hat{p}_k - p_k) + \lambda_2 \mathcal{L}_{Dice}(\hat{M}_{corner}, M_{corner})$$
Where $\hat{p}_k = (\hat{x}_k, \hat{y}_k)$ represents the normalized Cartesian coordinates for the four corners. Because regression heads alone can jitter across successive video frames, a Kalman state observer filters high-frequency frame tremor without inducing perceived temporal delay.
Rather than relying on iterative thin-plate spline meshes that drain battery reserves, Scan2PDF trains an encoder-decoder network to infer a dense 3D displacement vector field $\mathcal{D}(u, v) \rightarrow (x, y, z)$.
$$x(u, v) = u_0 + R \cdot \sin\left(\frac{u - u_0}{R}\right), \quad z(u, v) = R \cdot \left[1 - \cos\left(\frac{u - u_0}{R}\right)\right]$$
The model predicts local normal tangents $\mathbf{n}_i$, allowing Scan2PDF to perform backward grid unrolling in native C++ NDK, restoring curved lines to uniform horizontal alignments in under 0.90 seconds per page.
A photographed page $\mathbf{I}(x,y)$ in RGB color space is modeled as the element-wise product of surface reflectance $\mathbf{R}(x,y)$ and environmental illumination $\mathbf{L}(x,y)$:
$$\mathbf{I}(x,y) = \mathbf{R}(x,y) \odot \mathbf{L}(x,y)$$
Our 3.2 MB Auto Color network isolates the low-frequency illumination matrix $\hat{\mathbf{L}}(x,y)$ while preserving ink pigments inside $\mathbf{R}(x,y)$. Shadow-free normalization is then calculated directly:
$$\hat{\mathbf{R}}(x,y) = \min\left(255, \; \frac{\mathbf{I}(x,y) + \epsilon}{\hat{\mathbf{L}}(x,y) + \epsilon} \times \bar{\mathbf{L}}_{paper}\right)$$
Scan2PDF runs a lightweight U-Net segmenter generating a binary thumb mask $\mathbf{M}_{thumb} \in \{0, 1\}$. It enforces a strict border contiguity rule: any skin tone isolated within the document center is preserved (protecting passport photos and ID badges).
On held-out validation tests, the finger was successfully eliminated on 26 of 28 real phone photos, with false-positive triggers on only 2 of 104 fingerless pages.
$$q = \text{round}\left(\frac{r}{S}\right) + Z, \quad S = \frac{r_{max} - r_{min}}{q_{max} - q_{min}}$$
Quantizing weights from FP32 to signed INT8 shrank disk footprint by 74.2% (down to ~14MB aggregate) and sped up inference by 3.1x via vectorized ARM NEON SIMD acceleration.
Enterprise Partnership Custom Neural Vision & On-Device AI
Building low-latency, private, and quantized neural models requires specialized expertise in edge architecture, synthetic data simulation, and hardware acceleration. If your product requires custom computer vision, document intelligence, biometric segmentation, or offline neural OCR, Staksoft engineers custom on-device pipelines tailored specifically to your latency, size, and accuracy targets.
Edge CV & Detection
Ultra-lightweight object, keypoint, document boundary, and defect detection under 3MB.
Synthetic Datasets
Ray-traced photorealistic 3D synthetic training pipelines with deterministic ground-truth labels.
INT8/FP16 Quantization
Full post-training and quantization-aware training (QAT) optimized for Qualcomm NPU, Apple Neural Engine & Mali GPU.
Native C++/NDK Runtime
Zero-garbage-collection native Android & iOS engine integration with hardware delegate failovers.
Yes. Staksoft partners with enterprises, tech startups, and independent app publishers to design, train, compress, and deploy custom neural architectures directly into mobile apps. Whether you need real-time segmentation, offline OCR, or custom edge vision models, you can initiate a project discussion directly via our Contact Page.
No. All image scanning operations, boundary detection, spine flattening, Auto Color illumination normalization, finger erasure, and local AI OCR run entirely on the local device's CPU/GPU/NPU.
The models run on any device with Android 8.0 (API 26) or newer. The baseline latency benchmark of 12β25ms was measured on a Xiaomi Poco F1 equipped with a Qualcomm Snapdragon 845 SoC.
All weights are encrypted using AES-256 with dynamic memory decryption routines baked into compiled C++ binary layers. Decryption keys are derived dynamically from verified runtime APK signatures.
Test all four Document AI models directly on physical documents, thick open textbooks, and complex backgrounds. Download Scan2PDF for Android on Google Play today.
LLM integration, OCR, and on-device AI engineering from Staksoft.