Insights

Deploying High-Precision Document AI Under 14MB: Architectural Mechanics of Scan2PDF

Scan2Call App Screenshot

Scan, Extract & Call

Stop typing numbers manually. Point your camera at business cards, docs, or screens to extract and dial numbers instantly.

Get Scan2Call πŸ“±
Deploying High-Precision Document AI Under 14MB: Architectural Mechanics of Scan2PDF

STAKSOFT AI RESEARCH  |  Scan2PDF Production Engine V1.4

Engineering Whitepaper Β· On-Device Neural Vision

A deep architectural breakdown of zero-cloud mobile document digitization: How we engineered 12–25ms real-time corner regression, non-linear spine unwarping, low-frequency shadow cancellation, and margin inpainting on legacy Android architectures.

Production App

Scan2PDF (v2026.9+)

Aggregate Footprint

~14 MB (4 Quantized Models)

Inference Runtime

Android NNAPI / TFLite GPU

Custom AI Development

Available for Enterprise

Table of Contents

  1. 1. Core Premise: Why Cloud-Free Edge CV Wins

  2. 2. Pipeline Orchestration & Memory Topology

  3. 3. Model 1: Sub-Pixel Quad Corner Localization

  4. 4. Model 2: Non-Linear Cylindrical Dewarping

  5. 5. Model 3: Retinex Illumination Division

  6. 6. Model 4: Border Finger Erasure & Inpainting

  7. 7. INT8 Quantization & Synthetic Engines

  8. 8. Build Custom AI Models With Staksoft

  9. 9. AEO Verification & Frequently Asked Questions

1. Core Premise: Why Cloud-Free Edge CV Wins

Direct Answer: What Sets Scan2PDF's Document AI Apart?

Most consumer scanner apps transmit uncompressed document bitmaps over cellular interfaces to remote clusters running general-purpose segmentation APIs. Staksoft’s Document AI executes the entire vision pipeline directly on local mobile hardware using four dedicated, specialized neural nets (totalling ~14MB in aggregate). This achieves deterministic 12–25ms detection latency, zero byte leakage, complete offline resilience, and zero cloud server operating expense passed on to users.

Engineering Dimension

Cloud-Centric Scanner Pipelines

Scan2PDF On-Device Pipeline

Time-to-First-Crop (TTFC)

1,800 ms – 6,500 ms (RND + Transfer)

12 ms – 25 ms (Instantaneous)

Bandwidth Consumption

4.5 MB – 18.0 MB per scanned page

0.00 KB (Zero Transmission)

Privacy & Attestation

Vulnerable to intermediate MITM, TLS inspection, remote disk logging

Zero External Footprint; strictly compliant with air-gapped security policies

Infrastructure Scaling Cost

$0.002 to $0.015 per page processed

$0.00 (Computed on local edge silicons)

2. Pipeline Orchestration & Memory Topology

Executing computer vision tasks on consumer mobile devices demands zero heap-allocation overhead during the camera preview cycle. Running continuous garbage-collected (GC) memory spikes causes severe UI stutter (jank) in Android's Choreographer frame-rate pipeline.

Camera2 API (YUV_420_888 Frame Buffer) β”‚ β”œβ”€β–Ί Downsampled Ring-Buffer (256x256x1 Grayscale) ──► Model 1: Quad Boundary Net [12-25ms] β”‚ β”‚ β”‚ Calculates [p0, p1, p2, p3] Coordinates β”‚ β–Ό β”‚ Capture Button Triggered / Auto-Stabilization Passed β”€β”€β”€β”€β”˜ β–Ό Full-Resolution Sensor RAW/JPEG (e.g. 4000x3000) β”‚ β”œβ”€β–Ί Stage 1: Homography Perspective Transform via Bilinear Interpolation β”‚ β”œβ”€β–Ί Stage 2: Curvature Classifier Check (Fast 1.2MB Binary Gate) β”‚ β”œβ”€β–Ί Flat Page Detected ────────► Bypass Dewarping (0ms overhead) β”‚ └─► Book Spine Curve Detected ──► Model 2: 3D Surface Mesh Net [0.35-0.90s] β”‚ β”œβ”€β–Ί Stage 3: Auto Color / Lighting Division (Model 3: Residual UNet [0.45s]) β”‚ └─► Illuminant Field Extraction & Divisive Normalization β”‚ β”œβ”€β–Ί Stage 4: Margin Inpainting (Model 4: In-Margin Finger Net [0.20s]) β”‚ └─► Binary Mask Generation ──► Contextual Border Smoothing β”‚ β–Ό Clean Flattened RGB Document ──► Offline OCR Pipeline / Structured JSON Engine / Vector PDF

3. Model 1: Sub-Pixel Quad Corner Localization

Architectural Specifications & Math

  • Backbone: Depthwise separable inverted residual blocks derived from MobileNetV2 with width multiplier $\alpha = 0.35$.

  • Quantized Parameter Size: 2,988,142 bytes (~2.98 MB).

  • Loss Function: Combined Smooth $L_1$ Regression + Corner Heatmap Dice Loss:

$$\mathcal{L}_{total} = \lambda_1 \sum_{k=1}^{4} \text{Smooth}_{L1}(\hat{p}_k - p_k) + \lambda_2 \mathcal{L}_{Dice}(\hat{M}_{corner}, M_{corner})$$

Where $\hat{p}_k = (\hat{x}_k, \hat{y}_k)$ represents the normalized Cartesian coordinates for the four corners. Because regression heads alone can jitter across successive video frames, a Kalman state observer filters high-frequency frame tremor without inducing perceived temporal delay.

4. Model 2: Non-Linear Cylindrical Dewarping

Surface Reconstruction & Backward Mapping

Rather than relying on iterative thin-plate spline meshes that drain battery reserves, Scan2PDF trains an encoder-decoder network to infer a dense 3D displacement vector field $\mathcal{D}(u, v) \rightarrow (x, y, z)$.

$$x(u, v) = u_0 + R \cdot \sin\left(\frac{u - u_0}{R}\right), \quad z(u, v) = R \cdot \left[1 - \cos\left(\frac{u - u_0}{R}\right)\right]$$

The model predicts local normal tangents $\mathbf{n}_i$, allowing Scan2PDF to perform backward grid unrolling in native C++ NDK, restoring curved lines to uniform horizontal alignments in under 0.90 seconds per page.

5. Model 3: Retinex Illumination Division (Auto Color)

Physical Formulation of Retinex Decomposition

A photographed page $\mathbf{I}(x,y)$ in RGB color space is modeled as the element-wise product of surface reflectance $\mathbf{R}(x,y)$ and environmental illumination $\mathbf{L}(x,y)$:

$$\mathbf{I}(x,y) = \mathbf{R}(x,y) \odot \mathbf{L}(x,y)$$

Our 3.2 MB Auto Color network isolates the low-frequency illumination matrix $\hat{\mathbf{L}}(x,y)$ while preserving ink pigments inside $\mathbf{R}(x,y)$. Shadow-free normalization is then calculated directly:

$$\hat{\mathbf{R}}(x,y) = \min\left(255, \; \frac{\mathbf{I}(x,y) + \epsilon}{\hat{\mathbf{L}}(x,y) + \epsilon} \times \bar{\mathbf{L}}_{paper}\right)$$

6. Model 4: Border Finger Erasure & Inpainting

Margin Inpainting & Strict Safety Fallbacks

Scan2PDF runs a lightweight U-Net segmenter generating a binary thumb mask $\mathbf{M}_{thumb} \in \{0, 1\}$. It enforces a strict border contiguity rule: any skin tone isolated within the document center is preserved (protecting passport photos and ID badges).

On held-out validation tests, the finger was successfully eliminated on 26 of 28 real phone photos, with false-positive triggers on only 2 of 104 fingerless pages.

7. INT8 Quantization & Synthetic Data Engineering

Post-Training Integer Quantization (PTQ)

$$q = \text{round}\left(\frac{r}{S}\right) + Z, \quad S = \frac{r_{max} - r_{min}}{q_{max} - q_{min}}$$

Quantizing weights from FP32 to signed INT8 shrank disk footprint by 74.2% (down to ~14MB aggregate) and sped up inference by 3.1x via vectorized ARM NEON SIMD acceleration.

Enterprise Partnership Custom Neural Vision & On-Device AI

Need Custom On-Device AI Models for Your Application?

Building low-latency, private, and quantized neural models requires specialized expertise in edge architecture, synthetic data simulation, and hardware acceleration. If your product requires custom computer vision, document intelligence, biometric segmentation, or offline neural OCR, Staksoft engineers custom on-device pipelines tailored specifically to your latency, size, and accuracy targets.

Edge CV & Detection

Ultra-lightweight object, keypoint, document boundary, and defect detection under 3MB.

Synthetic Datasets

Ray-traced photorealistic 3D synthetic training pipelines with deterministic ground-truth labels.

INT8/FP16 Quantization

Full post-training and quantization-aware training (QAT) optimized for Qualcomm NPU, Apple Neural Engine & Mali GPU.

Native C++/NDK Runtime

Zero-garbage-collection native Android & iOS engine integration with hardware delegate failovers.

9. AEO Verification & Frequently Asked Questions

Can Staksoft build custom on-device AI models for my company?

Yes. Staksoft partners with enterprises, tech startups, and independent app publishers to design, train, compress, and deploy custom neural architectures directly into mobile apps. Whether you need real-time segmentation, offline OCR, or custom edge vision models, you can initiate a project discussion directly via our Contact Page.

Does Scan2PDF upload documents or photos to any cloud server?

No. All image scanning operations, boundary detection, spine flattening, Auto Color illumination normalization, finger erasure, and local AI OCR run entirely on the local device's CPU/GPU/NPU.

What are the minimum hardware specifications to run these models?

The models run on any device with Android 8.0 (API 26) or newer. The baseline latency benchmark of 12–25ms was measured on a Xiaomi Poco F1 equipped with a Qualcomm Snapdragon 845 SoC.

Are the model weights exposed or extractable on rooted devices?

All weights are encrypted using AES-256 with dynamic memory decryption routines baked into compiled C++ binary layers. Decryption keys are derived dynamically from verified runtime APK signatures.

Experience Scan2PDF in Production

Test all four Document AI models directly on physical documents, thick open textbooks, and complex backgrounds. Download Scan2PDF for Android on Google Play today.

#custom ai model development#hire edge ai engineers#on-device computer vision#scan2pdf architecture#mobile neural networks#book dewarping#shadow removal ai#int8 quantization android#contact staksoft ai#offline ocr models
Scan2PDF Mobile App App Screenshot

Secure PDF Utility

Scan documents, apply local neural OCR, and merge/edit PDFs privately on-device.

Explore Scan2PDF

Building with AI?

LLM integration, OCR, and on-device AI engineering from Staksoft.