# CUDA Distribution Problem - Complete Analysis ## Table of Contents 1. [Problem Overview](#problem-overview) 2. [Root Cause](#root-cause) 3. [Attempted Solutions](#attempted-solutions) 4. [Current Status](#current-status) 5. [Available Options](#available-options) 6. [Technical Details](#technical-details) 7. [Cost Analysis](#cost-analysis) 8. [Recommendations](#recommendations) --- ## Problem Overview ### Timeline of Issues **Original Problem (v0.1.0 - v0.1.11)** - Single server binary with CUDA support - Size: ~2.9GB - Issue: MSI installer build fails in GitHub Actions CI - Error: WiX Toolset cannot handle 3GB files efficiently **First Solution: Dual Binary System (v0.1.12)** - Split into CPU (295MB) and CUDA (2.37GB) binaries - CPU ships with installer - CUDA as optional download - Issue: GitHub Release assets have 2GB limit **Current Problem (Discovered during implementation)** - GitHub Release Asset Limit: **2GB hard maximum** - CUDA binary: **2.37GB** (370MB over limit) - Cannot upload to GitHub Releases --- ## Root Cause ### Why Is The CUDA Binary So Large? The size difference between CPU and CUDA builds: | Component | CPU Build | CUDA Build | Difference | |-----------|-----------|------------|------------| | PyTorch Core | ~150MB | ~150MB | - | | CPU Libraries (MKL/OpenBLAS) | ~100MB | - | -100MB | | CUDA Runtime | - | ~500MB | +500MB | | cuBLAS | - | ~350MB | +350MB | | cuDNN | - | ~1.2GB | +1.2GB | | NVRTC (CUDA Compiler) | - | ~90MB | +90MB | | Other CUDA libs | - | ~100MB | +100MB | | **Total** | **~295MB** | **~2.37GB** | **+2.07GB** | ### CUDA Dependencies Breakdown ``` torch/lib/ (CUDA build): ├── cudart64_12.dll (~0.5 MB) - CUDA Runtime ├── cublas64_12.dll (~100 MB) - Basic Linear Algebra ├── cublasLt64_12.dll (~200 MB) - Linear Algebra (optimized) ├── cudnn64_9.dll (~800 MB) - Deep Neural Networks ├── cudnn_*_infer64_9.dll (~400 MB) - DNN Inference ops ├── nvrtc64_*.dll (~50 MB) - Runtime Compiler ├── nvrtc-builtins64_*.dll (~40 MB) - Compiler builtins ├── torch_cuda.dll (~200 MB) - PyTorch CUDA bridge └── c10_cuda.dll (~20 MB) - Core CUDA utilities ``` **Why These Are Required:** - cuDNN is essential for neural network operations - cuBLAS handles all matrix operations (core of ML) - Cannot split or remove without breaking functionality --- ## Attempted Solutions ### Solution 1: Dual Binary System ✅ (Partially Successful) **Goal**: Split CPU and CUDA into separate downloads **Implementation**: ```bash # Build CPU-only (295MB) pip install torch --index-url https://download.pytorch.org/whl/cpu python build_binary.py cpu # Build CUDA (2.37GB) pip install torch --index-url https://download.pytorch.org/whl/cu121 python build_binary.py cuda ``` **Results**: - ✅ CPU binary: 295MB (fits in installer) - ✅ CI builds successfully - ✅ Installer size reduced from 3GB to ~500MB - ❌ CUDA binary still too large for GitHub **See**: `docs/dual-server-binaries.md` ### Solution 2: Compression Testing ❌ (Failed) **Goal**: Compress CUDA binary to fit under 2GB **Method**: 7z with maximum compression settings ```bash 7z a -t7z -m0=lzma2 -mx=9 -mfb=64 -md=32m -ms=on \ voicebox-server-cuda.7z voicebox-server-cuda.exe ``` **Results**: ``` Original: 2.37 GB (2,545,086,396 bytes) Compressed: 2.35 GB (2,519,381,264 bytes) Compression: 1.0% (only 24.5MB saved) GitHub Limit: 2.00 GB (2,147,483,648 bytes) Over by: 354.67 MB Status: FAILED - Still exceeds limit by 354MB ``` **Why Compression Failed**: - CUDA binaries are already optimized machine code - No redundant data to compress - Neural network kernels are highly compact - Libraries are already stripped of debug symbols **Conclusion**: Compression is not viable --- ## Current Status ### What Works - ✅ CPU binary builds successfully (295MB) - ✅ CUDA binary builds successfully (2.37GB) - ✅ Build scripts for both variants - ✅ CI workflow updated for dual binaries - ✅ Installer can be created with CPU binary ### What Doesn't Work - ❌ Cannot upload CUDA binary to GitHub Releases (exceeds 2GB limit) - ❌ Compression doesn't reduce size enough - ❌ No automated distribution path for CUDA binary ### Branch Status - Branch: `feat/dual-server-binaries` - Commits: Implementation complete - Testing: Local builds successful - Blocker: CUDA distribution path --- ## Available Options ### Option 1: AWS S3 Hosting (Recommended) **Description**: Host CUDA binary in Amazon S3 bucket **Pros**: - ✅ No file size limits (can handle multi-GB files) - ✅ Fast global CDN (CloudFront) - ✅ Reliable (99.99% uptime) - ✅ Pay only for usage - ✅ Easy CI integration - ✅ Version control (keep multiple releases) **Cons**: - ❌ Requires AWS account - ❌ Monthly costs (~$1-5/month) - ❌ Additional infrastructure to manage **Cost Estimate**: ``` Storage: 2.37 GB × $0.023/GB = $0.05/month Transfer: 100 downloads × 2.37GB × $0.09/GB = $21.33/month Total: ~$21-25/month for 100 downloads ~$2-5/month for 10-20 downloads ``` **Implementation**: ```yaml # .github/workflows/release.yml - name: Upload CUDA to S3 env: AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }} AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }} run: | aws s3 cp backend/cuda-release/voicebox-server-cuda-*.exe \ s3://voicebox-releases/cuda/${{ github.ref_name }}/ \ --acl public-read # Generate download URL echo "CUDA_URL=https://voicebox-releases.s3.amazonaws.com/cuda/${{ github.ref_name }}/voicebox-server-cuda-x86_64-pc-windows-msvc.exe" >> release_notes.txt ``` **User Experience**: 1. Install app normally (500MB installer) 2. App detects NVIDIA GPU 3. Shows: "Download CUDA support? (2.4GB)" 4. Downloads from S3: `https://voicebox-releases.s3.amazonaws.com/cuda/v0.1.12/voicebox-server-cuda.exe` 5. Saves to `%APPDATA%/voicebox/binaries/` 6. App restarts with CUDA server --- ### Option 2: Azure Blob Storage **Description**: Microsoft Azure alternative to S3 **Pros**: - ✅ Similar to S3 (no size limits, CDN, reliable) - ✅ Good if already using Azure - ✅ Competitive pricing - ✅ Global CDN with Azure CDN **Cons**: - ❌ Requires Azure account - ❌ Similar monthly costs - ❌ Less common in open source projects **Cost Estimate**: ``` Storage: $0.018/GB = $0.04/month Transfer: ~$20-25/month for 100 downloads ``` **Implementation**: ```yaml - name: Upload to Azure Blob env: AZURE_STORAGE_CONNECTION_STRING: ${{ secrets.AZURE_STORAGE }} run: | az storage blob upload \ --account-name voiceboxreleases \ --container-name cuda-binaries \ --name v${{ github.ref_name }}/voicebox-server-cuda.exe \ --file backend/cuda-release/voicebox-server-cuda-*.exe \ --tier Hot ``` --- ### Option 3: Cloudflare R2 **Description**: Cloudflare's S3-compatible object storage **Pros**: - ✅ S3-compatible API - ✅ **FREE egress (no bandwidth charges!)** - ✅ Cheaper than S3/Azure - ✅ Cloudflare CDN included - ✅ Good for open source projects **Cons**: - ❌ Requires Cloudflare account - ❌ Newer service (less mature than S3) **Cost Estimate**: ``` Storage: $0.015/GB = $0.04/month Egress: $0.00 (FREE!) Class A ops: Negligible Total: ~$0.04/month (essentially free!) ``` **Why This Is Attractive**: - Zero bandwidth costs (huge savings) - Perfect for open source distribution - S3-compatible (easy migration if needed) **Implementation**: Same as S3 (R2 is S3-compatible): ```yaml - name: Upload to R2 env: AWS_ACCESS_KEY_ID: ${{ secrets.R2_ACCESS_KEY_ID }} AWS_SECRET_ACCESS_KEY: ${{ secrets.R2_SECRET_ACCESS_KEY }} AWS_ENDPOINT_URL: https://.r2.cloudflarestorage.com run: | aws s3 cp backend/cuda-release/voicebox-server-cuda-*.exe \ s3://voicebox-releases/cuda/${{ github.ref_name }}/ \ --endpoint-url=$AWS_ENDPOINT_URL ``` --- ### Option 4: GitHub Packages (Container Registry) **Description**: Package CUDA binary as OCI/Docker artifact **Pros**: - ✅ Stays in GitHub ecosystem - ✅ No additional accounts needed - ✅ Free for public repos **Cons**: - ❌ Complex for desktop app distribution - ❌ Users need to extract from container - ❌ Awkward UX (not designed for binary distribution) - ❌ Requires Docker understanding **Not Recommended**: Containers aren't designed for desktop app binaries --- ### Option 5: Self-Hosted Server **Description**: Host on your own VPS/server **Pros**: - ✅ Full control - ✅ No cloud provider dependency - ✅ Predictable costs **Cons**: - ❌ Requires server maintenance - ❌ Bandwidth costs can be high - ❌ Uptime responsibility - ❌ Scaling challenges **Cost Estimate**: ``` VPS: $5-20/month (DigitalOcean, Linode) Bandwidth: $0.01-0.02/GB Total: $10-50/month depending on traffic ``` --- ### Option 6: Manual Distribution **Description**: Don't automate - provide manual download instructions **Pros**: - ✅ Zero cost - ✅ Zero infrastructure - ✅ Simple **Cons**: - ❌ Poor user experience - ❌ Manual upload to file host each release - ❌ Users must manually download and install - ❌ No automatic updates for CUDA binary - ❌ Increases support burden **Implementation**: ``` Release notes: "Windows users with NVIDIA GPUs can download CUDA support: 1. Download voicebox-server-cuda.exe from [Google Drive/Mega/etc] 2. Place in C:\Users\\AppData\Roaming\voicebox\binaries\ 3. Restart the app" ``` **Not Recommended**: Creates friction, support issues --- ### Option 7: Split CUDA Binary **Description**: Break CUDA binary into multiple <2GB chunks **Technical Approach**: ```python # Split binary split -b 2000M voicebox-server-cuda.exe cuda_part_ # Upload parts to GitHub (each <2GB) cuda_part_aa (2.0 GB) cuda_part_ab (0.37 GB) # App downloads and reassembles cat cuda_part_* > voicebox-server-cuda.exe ``` **Pros**: - ✅ Stays on GitHub - ✅ No external hosting **Cons**: - ❌ Complex download logic (multiple files) - ❌ Integrity checking required - ❌ More points of failure - ❌ Users must wait for multiple downloads - ❌ Still hacky solution **Complexity**: Medium-High --- ## Technical Details ### Current Build Output ``` backend/dist/ ├── voicebox-server.exe 295 MB (CPU-only) └── voicebox-server-cuda.exe 2.37 GB (CUDA) # After compression test: backend/dist/ └── voicebox-server-cuda.7z 2.35 GB (not viable) ``` ### CI Workflow Changes Required For external hosting (S3/R2/Azure): ```yaml # Current workflow (fails) - name: Upload CUDA server binary (Windows only) if: matrix.platform == 'windows-latest' uses: softprops/action-gh-release@v1 with: files: backend/cuda-release/voicebox-server-cuda-*.exe # ❌ Fails: >2GB draft: true # New workflow (S3 example) - name: Upload CUDA to S3 (Windows only) if: matrix.platform == 'windows-latest' env: AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }} AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }} run: | aws s3 cp backend/cuda-release/voicebox-server-cuda-*.exe \ s3://voicebox-releases/cuda/${{ github.ref_name }}/ \ --acl public-read # Generate release notes with download URL cat >> release_notes.md < { // Update UI: "Downloading CUDA support: 45% (1.2GB / 2.4GB)" }); // Verify checksum const checksum = await calculateChecksum(savePath); if (checksum !== EXPECTED_CHECKSUM) { throw new Error("Download corrupted"); } } ``` **Backend**: Already supports both binaries (no changes needed) --- ## Cost Analysis ### Monthly Cost Comparison (100 downloads/month) | Option | Storage | Bandwidth | Total/Month | Notes | |--------|---------|-----------|-------------|-------| | **Cloudflare R2** | $0.04 | $0.00 | **$0.04** | Best for open source | | AWS S3 | $0.05 | $21.33 | $21.38 | Good reliability | | Azure Blob | $0.04 | $20.00 | $20.04 | Azure ecosystem | | Self-hosted VPS | $10.00 | $2.37 | $12.37 | Maintenance overhead | | Manual | $0.00 | $0.00 | $0.00 | Poor UX | ### Annual Cost Comparison | Option | Year 1 | Year 2+ | Notes | |--------|--------|---------|-------| | **Cloudflare R2** | **$0.50** | **$0.50** | Essentially free | | AWS S3 | $256 | $256 | Predictable | | Self-hosted | $144 | $144 | Time cost | **Recommendation**: Cloudflare R2 (free egress = huge savings) --- ## Recommendations ### Recommended Solution: Cloudflare R2 **Why**: 1. **Cost**: Essentially free (~$0.04/month) 2. **Bandwidth**: Zero egress charges (unlimited downloads) 3. **CDN**: Cloudflare's global network included 4. **Compatibility**: S3-compatible API (easy to use) 5. **Perfect for open source**: No surprise bandwidth bills ### Implementation Priority **Phase 1: Setup (1-2 hours)** 1. Create Cloudflare R2 account 2. Create bucket: `voicebox-releases` 3. Generate API credentials 4. Add to GitHub Secrets **Phase 2: CI Integration (1-2 hours)** 1. Update `.github/workflows/release.yml` 2. Add R2 upload step 3. Generate release notes with download URL 4. Test with draft release **Phase 3: App Integration (4-6 hours)** 1. Add GPU detection on startup 2. Implement download manager UI 3. Add progress indicators 4. Implement checksum verification 5. Server restart logic **Phase 4: Documentation (1 hour)** 1. Update README with GPU instructions 2. Add troubleshooting guide 3. Document manual download process **Total Time**: ~8-12 hours of development ### Alternative: AWS S3 (If Already Using AWS) If you're already using AWS for other infrastructure, S3 is also a solid choice: - More mature than R2 - Extensive documentation - Familiar tooling - ~$20/month for moderate usage --- ## Open Questions 1. **Expected Download Volume**: How many CUDA downloads per month? - Affects cost calculations - Determines if R2's free egress is significant 2. **Update Strategy**: How to handle CUDA updates? - Option A: Version in URL path (keep all versions) - Option B: Overwrite latest (save space) 3. **Fallback Strategy**: What if cloud provider is down? - Mirror on multiple providers? - Graceful degradation to CPU? 4. **Telemetry**: Track CUDA download stats? - Helps with cost forecasting - User behavior insights --- ## Next Steps 1. **Research Phase** (You are here) - Evaluate cloud providers - Check terms of service - Test account creation 2. **Decision Phase** - Choose provider (Cloudflare R2 recommended) - Set up account - Configure billing alerts 3. **Implementation Phase** - Update CI workflow - Implement download manager - Test end-to-end flow 4. **Launch Phase** - Deploy to production - Monitor downloads - Gather user feedback --- ## References - **GitHub Release Limits**: https://docs.github.com/en/repositories/releasing-projects-on-github/about-releases - **Cloudflare R2 Pricing**: https://developers.cloudflare.com/r2/pricing/ - **AWS S3 Pricing**: https://aws.amazon.com/s3/pricing/ - **Compression Test Results**: `backend/test_cuda_compression.py` - **Dual Binary Implementation**: `docs/dual-server-binaries.md` --- ## Appendix: Alternative Approaches Considered ### A. Dynamic CUDA Loading **Idea**: Load CUDA DLLs dynamically at runtime **Why Not**: PyTorch requires CUDA DLLs at import time, can't lazy-load ### B. CUDA as Separate Package **Idea**: Python package with just CUDA libs **Why Not**: Still 2GB+, same problem ### C. Model Quantization **Idea**: Use smaller quantized models **Why Not**: Doesn't reduce CUDA runtime size ### D. Docker Distribution **Idea**: Distribute as Docker container **Why Not**: Poor fit for desktop app, requires Docker installed --- **Document Version**: 1.0 **Last Updated**: 2026-01-31 **Status**: Research Phase **Next Review**: After cloud provider decision