The short version
Yes, you can do all of this from your iPhone 16 Plus and iPad. The heavy work runs on computers you rent by the hour in the cloud. You use them through a web browser, and through the Linux box that already runs Chance AI.
- Base model: Qwen3.8‑27B. Its Apache 2.0 license lets you use it, change it, and sell things built on it. Gemma 4 31B (also Apache 2.0) is a good backup choice.
- Where it runs: RunPod, billed by the second. The minimum deposit is $10, with no subscription.
- First step: rent one 48 GB GPU for about 2 hours (A40 at $0.49/hr), load Qwen3.8‑27B, and plug it into Chance AI as a third provider called “My AI”. Expect about $1–$3 of your $10 deposit. Then delete the machine.
- Making it “yours”: train it on your own chats and code with LoRA (explained below). One round costs about $4–$20. Keep the results in a private place you control.
Quick terms. Model: the AI “brain” file. Open‑weight: you can download that file and keep it. GPU: the chip that runs AI. VRAM (GB): the GPU's memory; bigger models need more. LoRA: a small add‑on file that teaches a model your style or knowledge without retraining the whole thing. Endpoint: the web address your app sends questions to.
1 · Pick a base model you're allowed to own
You don't build the brain from nothing. You start from a free, open‑weight model whose license lets you modify it and use it commercially, then you shape it.
| Model family | License | Good for you? | Catch |
|---|---|---|---|
| Qwen3.8‑27B / Qwen3.6‑27B (Alibaba) | Apache 2.0 | Best pick. Strong at code, and runs on a single rented GPU. | Qwen's top “Max/Plus” models are only available through Alibaba's paid service; you can't download them. Check each model's license file. |
| Gemma 4 (Google, up to 31B) | Apache 2.0 (new with Gemma 4) | Great backup. Similar size. | Older Gemma 1–3 models use Google's stricter custom terms, so make sure the file you download says “Gemma 4”. |
| Mistral Small 4 | Apache 2.0 | Good, smaller option. | The larger Devstral 2 uses a “Modified MIT” license: companies earning more than $20M a month need a paid license. Some Mistral models aren't downloadable at all. |
| DeepSeek V4 Flash / Pro | MIT (very permissive) | Fine license, but too big to start with. | Flash has 284B parameters and Pro has 1.6T, so you'd need several large GPUs at once. Expensive. |
| Llama 4 (Meta) | Llama 4 Community License (not truly open source) | Usable, but has the most strings attached. | If you share it, you must show “Built with Llama” and start the model's name with “Llama”. You must follow Meta's use policy. Services with more than 700M users need Meta's permission. The multimodal (image‑reading) rights don't apply to people based in the EU. |
Plain meaning: under Apache 2.0 or MIT, the copy you download and the changes you make are yours to keep and use forever. Meta, Google, or Alibaba can't later turn off a file you already have. Your only duty is to keep their license notice with it.
2 · Renting GPU power from your phone (pay by the hour)
All of these run in a phone browser. “24 GB” is enough for small and medium models. “80 GB” is for bigger models and faster training.
| Host | 24 GB GPU | 80 GB GPU | Up‑front money | Storage fees / catches |
|---|---|---|---|---|
| RunPod (pods) | RTX 4090: $0.34/hr (community) · $0.74 (secure) RTX 3090: $0.22 / $0.50 | A100 80GB: $1.19–$1.39 (community) · $1.59 (secure) H100: $1.99–$2.89 | $10 minimum, prepaid and non‑refundable. You need at least 1 hour of credit to start a machine. | Disk: $0.10/GB/mo while running, but $0.20/GB/mo while stopped. Network storage: $0.07/GB/mo. Billed per second. 48 GB A40: $0.35 / $0.49. |
| Vast.ai (marketplace) | RTX 4090 from about $0.36/hr · RTX 3090 about $0.14 | A100 SXM4 from about $0.37 · H100 about $1.73 | $5 minimum deposit | Prices change constantly; these are the lowest live offers on Oct 7. The machines belong to many different hosts, so they're less private (filter for verified datacenters). A stopped machine still bills for storage until you destroy it. |
| Lambda | A10: $1.29/hr | H100 PCIe: $3.29/hr | No prepay: card on file, billed weekly by the minute | Storage: $0.20/GB/mo, billed even when nothing is using it. Pricier, but simple. |
| Modal | L4: $0.80/hr | A100 80GB: $2.50/hr · H100: $3.95/hr | $0 “Starter” plan that includes $30/mo of compute (not a subscription) | You set it up with code, so you'd do it from the Linux box, not by tapping. It shuts down when idle, so you pay nothing when it isn't in use. Storage: $0.09/GB/mo. |
| Hugging Face (Endpoints) | L4: $0.80/hr | A100 80GB: $2.50/hr · H100: $4.50/hr | Pay as you go (skip the PRO subscription) | Free accounts get 100 GB of private storage, which is a good place to keep your model files. |
| Together AI (fine‑tuning service) | Charges per amount of training text, not per hour (see Stage B). On‑demand H100 clusters cost $3.99/hr. | $5 minimum prepaid. Credits don't expire. | Easiest way to train without managing a machine, and you can download the result. | |
3 · The three stages, with real costs
Stage A: run an open model privately and add “My AI” to Chance AI
- On RunPod, start a pod with a ready‑made vLLM or Ollama template. These are free programs that run a model and give it an OpenAI‑style web address. Use a 48 GB A40 ($0.49/hr) or a 24 GB 4090 ($0.34–$0.74/hr).
- Load Qwen3.8‑27B. On a 24 GB card, use a 4‑bit (compressed) copy. Set a password (vLLM's
--api-key). - In Chance AI, add a third mode, “My AI”, next to Standard and Private. It points at that address. Your app's Private mode already talks to Tinfoil and Venice in this same “OpenAI‑style” format, so this is a small change the App Builder can make.
Cost: about $1–$3 for a 2–4 hour test. Watch out: leaving it on 24/7 costs about $245–$533 a month ($0.34–$0.74 × 720 hours). So turn it on only when you need it, or use RunPod “Serverless” (24 GB at $0.69/hr, charged only while it's answering, with a slow first reply).
Stage B: customize it with LoRA on your own data and code
- Easiest from a phone (Together AI): upload one file of example conversations and pay per token (about ¾ of a word). Qwen 27B‑class costs $1.05 per million tokens, with a $4 minimum per job. Example: 2 million tokens of your data, read 3 times (3 “epochs”) = 6M tokens ≈ $6.30. Then download your LoRA file.
- Do‑it‑yourself (RunPod plus the free Unsloth tool): training Qwen 27B with LoRA needs at least about 22 GB of GPU memory, so rent the 48 GB A40 for 2–6 hours ≈ $1–$3 (plus setup time).
- Your LoRA file is small, usually megabytes to a few GB. It belongs to you and works on top of the free base model.
Stage C: what “creating your own AI” really means long term
- Realistic and yours: your own code (Chance AI), plus your own data, plus your own fine‑tuned version of an open model. Repeat it every few months as better base models come out. About $5–$100 per round. This is how most companies “make their own AI”.
- Training from scratch, tiny version (for learning): Karpathy's free nanochat project trains a small chatbot from zero on 8 rented H100s in about 2–4 hours, for about $48–$100. It works, but it's roughly at GPT‑2 level: a fun experiment, not a daily assistant.
- Training from scratch, competitive with today's best: DeepSeek‑V3's final training run alone cost about $5.6 million (2.79M GPU‑hours at $2), not counting research. That's not a sensible goal for one person. Fine‑tuning gets you 95% of the ownership for 0.001% of the cost.
4 · Bringing all your Chance AI code over
What's in it now (from a read‑only look at the folder): the Flask server (server.py), agents (agents.py), the App Builder (builder.py), Private‑mode providers (private_providers.py), the cloud‑computer feature (computer.py), extras, voice, the web app (index.html, app.js, app.css), tests, and a data/ folder (settings, agents and their history, analytics). There is also a Replit copy.
- Put it in a private GitHub repo. Free accounts get unlimited private repos. Keep passwords and API keys out of the repo:
data/already has a.gitignore(a list of files git skips), so add your key files to it. - Run it anywhere. Keep it on the always‑on Linux box it uses today (cheap, already working with your Cloudflare tunnel), or run it on the same rented GPU machine as your model.
- Add “My AI” as a third provider next to Standard (Gemini/OpenRouter) and Private (Tinfoil/Venice). It points at your model's OpenAI‑style address from vLLM or Ollama.
- Outside services keep working. Gemini, OpenRouter, Inworld voice, and Tinfoil keep running on your own keys. You can replace them one at a time with your own model whenever you like.
- Your app becomes your training data. Your Chance AI chats, agent histories, playbooks, and change requests can be turned into the example conversations Stage B needs. Remove anything private first. Catch: your own messages are fine to use, but some providers limit training on their AI's replies. Google's Gemini API terms say you may not use it “to develop models that compete with” Gemini. Train on your own writing, on replies from open models, or on replies you've corrected yourself.
5 · How your code and model stay yours
- Code: private GitHub repo (free). Also download a ZIP of it now and then to iCloud Drive or the Files app.
- Model files: save your LoRA file, and if you want the “merged” full model, to storage you control. Options: a private Hugging Face repo (100 GB free), RunPod network storage ($0.07/GB/mo), or the Linux box.
- Exporting:
- Together: run
together fine-tuning download <job-id> --checkpoint-type adapter(ormerged) from the Linux box. - RunPod or Vast: before you delete the machine, upload the output folder to your private Hugging Face repo or copy it to the Linux box.
- Together: run
- Rule of thumb: if a file exists only on a rented machine, it isn't safe yet. Once it's in two places you control, it's yours.
6 · Honest note
Rented cloud GPUs are far more powerful than any home computer you'd buy. You can rent a top data‑center chip (an H100) for about $1.99–$3.99/hr, and nothing you could put in a home matches it. But: it needs an internet connection, you pay every minute a machine is on (even if you forget it), and stopped machines still charge for storage. Always delete machines when you're done, and keep deposits small.
Recommended starter path (all from iPhone/iPad)
- Today, free: create a private GitHub repo and put the Chance AI code in it, with keys excluded.
- First paid step, about $10 deposit (about $1–$3 used): on RunPod, start a 48 GB A40 with the vLLM template, load Qwen3.8‑27B, and try it from Chance AI as “My AI”. Then delete the pod.
- Next, about $5–$10 more: export about 1–2M tokens of your own cleaned Chance AI conversations. Run one LoRA fine‑tune (Together, $5 minimum, or on RunPod). Download the result to your private Hugging Face repo.
- Then: load your LoRA on top of Qwen in “My AI”. That's your AI: your code, your data, your model files.
Total to get started: about $15–$20, all pay‑as‑you‑go, with no subscriptions.
Sources
- RunPod pricing (updated Sept 27, 2026): runpod.io/pricing · billing and $10 minimum: docs.runpod.io/accounts-billing/billing
- Vast.ai: vast.ai/pricing · live lows from GPU Radar (tracks Vast marketplace, Oct 7, 2026) · $5 minimum and storage billing: docs.vast.ai quickstart, billing
- Lambda: lambda.ai/pricing · docs.lambda.ai billing
- Modal: modal.com/pricing
- Hugging Face: huggingface.co/pricing · storage limits
- Together AI: together.ai/pricing · credits ($5 minimum) · fine-tune download CLI
- Licenses: Qwen3.8‑27B · Qwen3.6‑27B license · Gemma 4 Apache 2.0 · older Gemma terms · DeepSeek V4 (MIT) · Mistral Small 4 · Devstral 2 · Llama 4 license · Llama 4 use policy
- Tools: vLLM OpenAI‑compatible server · Ollama OpenAI compatibility · Unsloth VRAM requirements
- Cost scale: nanochat (~$100 from scratch) · DeepSeek‑V3 report ($5.576M final run)
- Training‑data catch: Gemini API terms · Private repos: github.com/pricing
Prices change often, especially on Vast.ai, so check the live page before you rent. On RunPod, “Community” means cheaper machines from vetted outside hosts, and “Secure” means higher‑reliability data centers.