- Documents entire problem from 3GB installer to GitHub 2GB limit - Analyzes compression test failure (1% reduction) - Compares 7 different hosting options - Cost analysis for each approach - Recommends Cloudflare R2 (free egress, ~/usr/bin/bash.04/month) - Technical implementation details for all options - Complete research document for decision making
16 KiB
CUDA Distribution Problem - Complete Analysis
Table of Contents
- Problem Overview
- Root Cause
- Attempted Solutions
- Current Status
- Available Options
- Technical Details
- Cost Analysis
- Recommendations
Problem Overview
Timeline of Issues
Original Problem (v0.1.0 - v0.1.11)
- Single server binary with CUDA support
- Size: ~2.9GB
- Issue: MSI installer build fails in GitHub Actions CI
- Error: WiX Toolset cannot handle 3GB files efficiently
First Solution: Dual Binary System (v0.1.12)
- Split into CPU (295MB) and CUDA (2.37GB) binaries
- CPU ships with installer
- CUDA as optional download
- Issue: GitHub Release assets have 2GB limit
Current Problem (Discovered during implementation)
- GitHub Release Asset Limit: 2GB hard maximum
- CUDA binary: 2.37GB (370MB over limit)
- Cannot upload to GitHub Releases
Root Cause
Why Is The CUDA Binary So Large?
The size difference between CPU and CUDA builds:
| Component | CPU Build | CUDA Build | Difference |
|---|---|---|---|
| PyTorch Core | ~150MB | ~150MB | - |
| CPU Libraries (MKL/OpenBLAS) | ~100MB | - | -100MB |
| CUDA Runtime | - | ~500MB | +500MB |
| cuBLAS | - | ~350MB | +350MB |
| cuDNN | - | ~1.2GB | +1.2GB |
| NVRTC (CUDA Compiler) | - | ~90MB | +90MB |
| Other CUDA libs | - | ~100MB | +100MB |
| Total | ~295MB | ~2.37GB | +2.07GB |
CUDA Dependencies Breakdown
torch/lib/ (CUDA build):
├── cudart64_12.dll (~0.5 MB) - CUDA Runtime
├── cublas64_12.dll (~100 MB) - Basic Linear Algebra
├── cublasLt64_12.dll (~200 MB) - Linear Algebra (optimized)
├── cudnn64_9.dll (~800 MB) - Deep Neural Networks
├── cudnn_*_infer64_9.dll (~400 MB) - DNN Inference ops
├── nvrtc64_*.dll (~50 MB) - Runtime Compiler
├── nvrtc-builtins64_*.dll (~40 MB) - Compiler builtins
├── torch_cuda.dll (~200 MB) - PyTorch CUDA bridge
└── c10_cuda.dll (~20 MB) - Core CUDA utilities
Why These Are Required:
- cuDNN is essential for neural network operations
- cuBLAS handles all matrix operations (core of ML)
- Cannot split or remove without breaking functionality
Attempted Solutions
Solution 1: Dual Binary System ✅ (Partially Successful)
Goal: Split CPU and CUDA into separate downloads
Implementation:
# Build CPU-only (295MB)
pip install torch --index-url https://download.pytorch.org/whl/cpu
python build_binary.py cpu
# Build CUDA (2.37GB)
pip install torch --index-url https://download.pytorch.org/whl/cu121
python build_binary.py cuda
Results:
- ✅ CPU binary: 295MB (fits in installer)
- ✅ CI builds successfully
- ✅ Installer size reduced from 3GB to ~500MB
- ❌ CUDA binary still too large for GitHub
See: docs/dual-server-binaries.md
Solution 2: Compression Testing ❌ (Failed)
Goal: Compress CUDA binary to fit under 2GB
Method: 7z with maximum compression settings
7z a -t7z -m0=lzma2 -mx=9 -mfb=64 -md=32m -ms=on \
voicebox-server-cuda.7z voicebox-server-cuda.exe
Results:
Original: 2.37 GB (2,545,086,396 bytes)
Compressed: 2.35 GB (2,519,381,264 bytes)
Compression: 1.0% (only 24.5MB saved)
GitHub Limit: 2.00 GB (2,147,483,648 bytes)
Over by: 354.67 MB
Status: FAILED - Still exceeds limit by 354MB
Why Compression Failed:
- CUDA binaries are already optimized machine code
- No redundant data to compress
- Neural network kernels are highly compact
- Libraries are already stripped of debug symbols
Conclusion: Compression is not viable
Current Status
What Works
- ✅ CPU binary builds successfully (295MB)
- ✅ CUDA binary builds successfully (2.37GB)
- ✅ Build scripts for both variants
- ✅ CI workflow updated for dual binaries
- ✅ Installer can be created with CPU binary
What Doesn't Work
- ❌ Cannot upload CUDA binary to GitHub Releases (exceeds 2GB limit)
- ❌ Compression doesn't reduce size enough
- ❌ No automated distribution path for CUDA binary
Branch Status
- Branch:
feat/dual-server-binaries - Commits: Implementation complete
- Testing: Local builds successful
- Blocker: CUDA distribution path
Available Options
Option 1: AWS S3 Hosting (Recommended)
Description: Host CUDA binary in Amazon S3 bucket
Pros:
- ✅ No file size limits (can handle multi-GB files)
- ✅ Fast global CDN (CloudFront)
- ✅ Reliable (99.99% uptime)
- ✅ Pay only for usage
- ✅ Easy CI integration
- ✅ Version control (keep multiple releases)
Cons:
- ❌ Requires AWS account
- ❌ Monthly costs (~$1-5/month)
- ❌ Additional infrastructure to manage
Cost Estimate:
Storage: 2.37 GB × $0.023/GB = $0.05/month
Transfer: 100 downloads × 2.37GB × $0.09/GB = $21.33/month
Total: ~$21-25/month for 100 downloads
~$2-5/month for 10-20 downloads
Implementation:
# .github/workflows/release.yml
- name: Upload CUDA to S3
env:
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
run: |
aws s3 cp backend/cuda-release/voicebox-server-cuda-*.exe \
s3://voicebox-releases/cuda/${{ github.ref_name }}/ \
--acl public-read
# Generate download URL
echo "CUDA_URL=https://voicebox-releases.s3.amazonaws.com/cuda/${{ github.ref_name }}/voicebox-server-cuda-x86_64-pc-windows-msvc.exe" >> release_notes.txt
User Experience:
- Install app normally (500MB installer)
- App detects NVIDIA GPU
- Shows: "Download CUDA support? (2.4GB)"
- Downloads from S3:
https://voicebox-releases.s3.amazonaws.com/cuda/v0.1.12/voicebox-server-cuda.exe - Saves to
%APPDATA%/voicebox/binaries/ - App restarts with CUDA server
Option 2: Azure Blob Storage
Description: Microsoft Azure alternative to S3
Pros:
- ✅ Similar to S3 (no size limits, CDN, reliable)
- ✅ Good if already using Azure
- ✅ Competitive pricing
- ✅ Global CDN with Azure CDN
Cons:
- ❌ Requires Azure account
- ❌ Similar monthly costs
- ❌ Less common in open source projects
Cost Estimate:
Storage: $0.018/GB = $0.04/month
Transfer: ~$20-25/month for 100 downloads
Implementation:
- name: Upload to Azure Blob
env:
AZURE_STORAGE_CONNECTION_STRING: ${{ secrets.AZURE_STORAGE }}
run: |
az storage blob upload \
--account-name voiceboxreleases \
--container-name cuda-binaries \
--name v${{ github.ref_name }}/voicebox-server-cuda.exe \
--file backend/cuda-release/voicebox-server-cuda-*.exe \
--tier Hot
Option 3: Cloudflare R2
Description: Cloudflare's S3-compatible object storage
Pros:
- ✅ S3-compatible API
- ✅ FREE egress (no bandwidth charges!)
- ✅ Cheaper than S3/Azure
- ✅ Cloudflare CDN included
- ✅ Good for open source projects
Cons:
- ❌ Requires Cloudflare account
- ❌ Newer service (less mature than S3)
Cost Estimate:
Storage: $0.015/GB = $0.04/month
Egress: $0.00 (FREE!)
Class A ops: Negligible
Total: ~$0.04/month (essentially free!)
Why This Is Attractive:
- Zero bandwidth costs (huge savings)
- Perfect for open source distribution
- S3-compatible (easy migration if needed)
Implementation: Same as S3 (R2 is S3-compatible):
- name: Upload to R2
env:
AWS_ACCESS_KEY_ID: ${{ secrets.R2_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.R2_SECRET_ACCESS_KEY }}
AWS_ENDPOINT_URL: https://<account-id>.r2.cloudflarestorage.com
run: |
aws s3 cp backend/cuda-release/voicebox-server-cuda-*.exe \
s3://voicebox-releases/cuda/${{ github.ref_name }}/ \
--endpoint-url=$AWS_ENDPOINT_URL
Option 4: GitHub Packages (Container Registry)
Description: Package CUDA binary as OCI/Docker artifact
Pros:
- ✅ Stays in GitHub ecosystem
- ✅ No additional accounts needed
- ✅ Free for public repos
Cons:
- ❌ Complex for desktop app distribution
- ❌ Users need to extract from container
- ❌ Awkward UX (not designed for binary distribution)
- ❌ Requires Docker understanding
Not Recommended: Containers aren't designed for desktop app binaries
Option 5: Self-Hosted Server
Description: Host on your own VPS/server
Pros:
- ✅ Full control
- ✅ No cloud provider dependency
- ✅ Predictable costs
Cons:
- ❌ Requires server maintenance
- ❌ Bandwidth costs can be high
- ❌ Uptime responsibility
- ❌ Scaling challenges
Cost Estimate:
VPS: $5-20/month (DigitalOcean, Linode)
Bandwidth: $0.01-0.02/GB
Total: $10-50/month depending on traffic
Option 6: Manual Distribution
Description: Don't automate - provide manual download instructions
Pros:
- ✅ Zero cost
- ✅ Zero infrastructure
- ✅ Simple
Cons:
- ❌ Poor user experience
- ❌ Manual upload to file host each release
- ❌ Users must manually download and install
- ❌ No automatic updates for CUDA binary
- ❌ Increases support burden
Implementation:
Release notes:
"Windows users with NVIDIA GPUs can download CUDA support:
1. Download voicebox-server-cuda.exe from [Google Drive/Mega/etc]
2. Place in C:\Users\<YourName>\AppData\Roaming\voicebox\binaries\
3. Restart the app"
Not Recommended: Creates friction, support issues
Option 7: Split CUDA Binary
Description: Break CUDA binary into multiple <2GB chunks
Technical Approach:
# Split binary
split -b 2000M voicebox-server-cuda.exe cuda_part_
# Upload parts to GitHub (each <2GB)
cuda_part_aa (2.0 GB)
cuda_part_ab (0.37 GB)
# App downloads and reassembles
cat cuda_part_* > voicebox-server-cuda.exe
Pros:
- ✅ Stays on GitHub
- ✅ No external hosting
Cons:
- ❌ Complex download logic (multiple files)
- ❌ Integrity checking required
- ❌ More points of failure
- ❌ Users must wait for multiple downloads
- ❌ Still hacky solution
Complexity: Medium-High
Technical Details
Current Build Output
backend/dist/
├── voicebox-server.exe 295 MB (CPU-only)
└── voicebox-server-cuda.exe 2.37 GB (CUDA)
# After compression test:
backend/dist/
└── voicebox-server-cuda.7z 2.35 GB (not viable)
CI Workflow Changes Required
For external hosting (S3/R2/Azure):
# Current workflow (fails)
- name: Upload CUDA server binary (Windows only)
if: matrix.platform == 'windows-latest'
uses: softprops/action-gh-release@v1
with:
files: backend/cuda-release/voicebox-server-cuda-*.exe # ❌ Fails: >2GB
draft: true
# New workflow (S3 example)
- name: Upload CUDA to S3 (Windows only)
if: matrix.platform == 'windows-latest'
env:
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
run: |
aws s3 cp backend/cuda-release/voicebox-server-cuda-*.exe \
s3://voicebox-releases/cuda/${{ github.ref_name }}/ \
--acl public-read
# Generate release notes with download URL
cat >> release_notes.md <<EOF
### GPU Acceleration (Windows)
Download CUDA support for NVIDIA GPUs:
[voicebox-server-cuda.exe](https://voicebox-releases.s3.amazonaws.com/cuda/${{ github.ref_name }}/voicebox-server-cuda-x86_64-pc-windows-msvc.exe)
Size: 2.37 GB
EOF
App Changes Required
Frontend (Tauri): Download manager
// src/lib/cuda-downloader.ts
const CUDA_DOWNLOAD_URL =
"https://voicebox-releases.s3.amazonaws.com/cuda/v{VERSION}/voicebox-server-cuda.exe";
async function downloadCudaBinary(version: string) {
const url = CUDA_DOWNLOAD_URL.replace("{VERSION}", version);
const savePath = path.join(app.getPath("userData"), "binaries", "voicebox-server-cuda.exe");
// Download with progress
await downloadFile(url, savePath, (progress) => {
// Update UI: "Downloading CUDA support: 45% (1.2GB / 2.4GB)"
});
// Verify checksum
const checksum = await calculateChecksum(savePath);
if (checksum !== EXPECTED_CHECKSUM) {
throw new Error("Download corrupted");
}
}
Backend: Already supports both binaries (no changes needed)
Cost Analysis
Monthly Cost Comparison (100 downloads/month)
| Option | Storage | Bandwidth | Total/Month | Notes |
|---|---|---|---|---|
| Cloudflare R2 | $0.04 | $0.00 | $0.04 | Best for open source |
| AWS S3 | $0.05 | $21.33 | $21.38 | Good reliability |
| Azure Blob | $0.04 | $20.00 | $20.04 | Azure ecosystem |
| Self-hosted VPS | $10.00 | $2.37 | $12.37 | Maintenance overhead |
| Manual | $0.00 | $0.00 | $0.00 | Poor UX |
Annual Cost Comparison
| Option | Year 1 | Year 2+ | Notes |
|---|---|---|---|
| Cloudflare R2 | $0.50 | $0.50 | Essentially free |
| AWS S3 | $256 | $256 | Predictable |
| Self-hosted | $144 | $144 | Time cost |
Recommendation: Cloudflare R2 (free egress = huge savings)
Recommendations
Recommended Solution: Cloudflare R2
Why:
- Cost: Essentially free (~$0.04/month)
- Bandwidth: Zero egress charges (unlimited downloads)
- CDN: Cloudflare's global network included
- Compatibility: S3-compatible API (easy to use)
- Perfect for open source: No surprise bandwidth bills
Implementation Priority
Phase 1: Setup (1-2 hours)
- Create Cloudflare R2 account
- Create bucket:
voicebox-releases - Generate API credentials
- Add to GitHub Secrets
Phase 2: CI Integration (1-2 hours)
- Update
.github/workflows/release.yml - Add R2 upload step
- Generate release notes with download URL
- Test with draft release
Phase 3: App Integration (4-6 hours)
- Add GPU detection on startup
- Implement download manager UI
- Add progress indicators
- Implement checksum verification
- Server restart logic
Phase 4: Documentation (1 hour)
- Update README with GPU instructions
- Add troubleshooting guide
- Document manual download process
Total Time: ~8-12 hours of development
Alternative: AWS S3 (If Already Using AWS)
If you're already using AWS for other infrastructure, S3 is also a solid choice:
- More mature than R2
- Extensive documentation
- Familiar tooling
- ~$20/month for moderate usage
Open Questions
-
Expected Download Volume: How many CUDA downloads per month?
- Affects cost calculations
- Determines if R2's free egress is significant
-
Update Strategy: How to handle CUDA updates?
- Option A: Version in URL path (keep all versions)
- Option B: Overwrite latest (save space)
-
Fallback Strategy: What if cloud provider is down?
- Mirror on multiple providers?
- Graceful degradation to CPU?
-
Telemetry: Track CUDA download stats?
- Helps with cost forecasting
- User behavior insights
Next Steps
-
Research Phase (You are here)
- Evaluate cloud providers
- Check terms of service
- Test account creation
-
Decision Phase
- Choose provider (Cloudflare R2 recommended)
- Set up account
- Configure billing alerts
-
Implementation Phase
- Update CI workflow
- Implement download manager
- Test end-to-end flow
-
Launch Phase
- Deploy to production
- Monitor downloads
- Gather user feedback
References
- GitHub Release Limits: https://docs.github.com/en/repositories/releasing-projects-on-github/about-releases
- Cloudflare R2 Pricing: https://developers.cloudflare.com/r2/pricing/
- AWS S3 Pricing: https://aws.amazon.com/s3/pricing/
- Compression Test Results:
backend/test_cuda_compression.py - Dual Binary Implementation:
docs/dual-server-binaries.md
Appendix: Alternative Approaches Considered
A. Dynamic CUDA Loading
Idea: Load CUDA DLLs dynamically at runtime Why Not: PyTorch requires CUDA DLLs at import time, can't lazy-load
B. CUDA as Separate Package
Idea: Python package with just CUDA libs Why Not: Still 2GB+, same problem
C. Model Quantization
Idea: Use smaller quantized models Why Not: Doesn't reduce CUDA runtime size
D. Docker Distribution
Idea: Distribute as Docker container Why Not: Poor fit for desktop app, requires Docker installed
Document Version: 1.0 Last Updated: 2026-01-31 Status: Research Phase Next Review: After cloud provider decision