← Back to Blog

Running a small team's internal app on Cloud Run: what we set up so it stays cheap and stays up

  • Google Cloud
Running a small team's internal app on Cloud Run: what we set up so it stays cheap and stays up

Say a 14-person logistics company has one internal app: drivers mark deliveries done on a phone, the office watches the list fill in. It was built in a week, it runs in a container, and somebody deployed it to Cloud Run because that was the shortest path from "it works on my machine" to "it works in a parking lot". Eight months later two things have happened. The bill is not what anyone expected, and nobody is certain who can open the URL.

Neither is bad luck. Cloud Run is a good home for an app like this, and its defaults are sensible for a public website with real traffic. An internal app used by fourteen people is a different shape, and about five settings decide whether it costs a few dollars a month or a few hundred. This is the order we check them in on a client service. Every figure below comes from Google's own documentation, checked 5 October 2026.

What a Cloud Run service actually bills you for

Cloud Run has two billing settings, and the difference between them is the biggest single number on the invoice.

The default is request-based billing. Instances are charged while they process a request, plus the time they spend starting up and shutting down. An instance that sits idle and is not a minimum instance is not charged at all, and Google's documentation says an instance will never stay idle for more than 15 minutes after processing a request unless minimum instances keep it alive. For an app that is busy from 7am to 6pm and silent the rest of the time, that is close to ideal.

The other setting is instance-based billing, previously called "CPU always allocated". It bills the entire lifetime of every instance, from start to termination, with a minimum of one minute, whether or not a request arrives. It earns that cost when your container does work between requests: a background queue, a scheduler running inside the process, a cache it warms on its own. A form that saves a row does not need it.

The published rates for request-based billing are USD 0.000024 per vCPU-second and USD 0.0000025 per GiB-second of active time, plus USD 0.40 per million requests, list price checked 5 October 2026. Rates differ by region, so read the row for the region you are actually deploying to. There is also a monthly free tier: the first 180,000 vCPU-seconds, 360,000 GiB-seconds and 2 million requests. A fourteen-person internal app that scales to zero overnight often lives inside that free tier, and its invoice line is rounding error.

The two settings that quietly create a bill

When a small-business Cloud Run bill is surprising, it is almost always one of these two.

The first is minimum instances. Setting it above zero keeps instances warm so the first request of the morning is not slow. That is a real problem worth solving, and it is not free: under request-based billing, minimum instances are charged at a reduced idle rate for CPU and memory while they wait, and that rate runs 24 hours a day, including the sixteen hours nobody is using the app. One warm instance is a defensible choice. One warm instance with instance-based billing switched on as well is a small server you are renting around the clock without meaning to.

The second is instance-based billing left on out of habit, usually because somebody once had a background task and nobody switched it back afterwards. Check which mode your service is on before you tune anything else, because every other number on this page is read differently under each one.

If you do want warm instances, the honest version is one, in the region your staff are in, under request-based billing, and only after you have measured how slow a cold start really is for your container. A container that starts in 400ms does not need to be kept awake.

The three request settings nobody revisits

Concurrency is how many requests one instance handles at the same time. A service deployed from the Google Cloud console defaults to 80. Deployed with the gcloud CLI or Terraform, the default is 80 times the number of vCPUs, and it drops to 1 if you requested less than one vCPU. The ceiling is 1,000. For an internal app the default is almost always right, and the trap is the fractional CPU: ask for half a vCPU to save money and you get a service that handles one request per instance, so fourteen people clicking at the same time spin up fourteen instances.

Maximum instances defaults to 100 per revision. That default exists for a website that might get linked from somewhere popular. For an app with fourteen known users, 100 is not a safety limit, it is a blank cheque written against your own bugs: a retry loop in the phone client, or a crawler that found an endpoint you forgot to close, can scale you to a hundred instances, and the only thing that notices is the invoice. We set a cap that reflects the real user count, usually a single digit, and treat hitting it as an alert rather than as something to raise.

Request timeout defaults to 300 seconds and can be set as high as 3,600. Five minutes is a long time to hold a connection open for an app whose slowest page is a report. Lowering it means a stuck request fails fast instead of billing CPU while it hangs. Google's documentation notes that beyond 15 minutes you should implement retries and expect clients to reconnect, which is a good hint that long timeouts belong to a different kind of workload.

Who can call the URL, and where the password lives

A Cloud Run service is reachable over HTTPS by any identity that holds the Cloud Run Invoker role (roles/run.invoker) on it. Granting that role to the special principal allUsers is what makes a service public. On an internal app, that grant should not exist. Grant the role to the people who need it instead:

gcloud run services add-iam-policy-binding internal-app \
  --region=REGION \
  --member=group:operations@yourcompany.com \
  --role=roles/run.invoker

Deploy with --no-allow-unauthenticated and the service rejects unauthenticated requests with an HTTP 403, so a leaked URL stops being an incident. One thing to know before you tidy this up: turning the Invoker IAM check back on in the console removes any existing allUsers binding for you, which is convenient and also means anything you did not know was calling the service anonymously will stop working that minute. Find those callers first.

Secrets are the other half of the same job. Cloud Run reads Secret Manager secrets either as an environment variable or as a file mounted into the container, both through --set-secrets, where a key starting with a slash is a mount path and anything else is an environment variable:

gcloud run deploy internal-app \
  --region=REGION \
  --set-secrets=/secrets/db/password=db-password:latest \
  --no-allow-unauthenticated

The two forms fail differently, and the difference is worth choosing on purpose. An environment variable secret is fetched before the instance starts, so if the secret is unreachable the instance does not start and you learn about it at deploy time. A mounted secret is not checked at startup; the read fails later, at runtime, inside whatever code path happens to need it. We mount the secrets that rotate, and use environment variables for the ones the app cannot function without.

Logs: the 30-day default and what happens after it

Everything the container writes to standard output lands in Cloud Logging, in the _Default bucket, which keeps logs for 30 days. You can set that anywhere from 1 day to 3,650 days. Two facts make the decision simple. Retention beyond the default 30 days is billed at USD 0.01 per GiB per month, list price checked 5 October 2026. And shortening retention to less than 30 days saves nothing, because the first 30 days are included with ingestion.

So the only retention change that moves money is a long one, and a chatty internal app with debug logging left switched on is the service most likely to make it expensive. Turn the log level down before you turn the retention up.

Cloud Run or a small VM

The comparison people expect to be about price is mostly about who patches the operating system. A small VM running the same container has a predictable monthly cost and an endless list of small obligations: the OS reaches end of life, the disk fills with logs, a certificate expires, somebody has to be the person who logs in. Cloud Run removes that list, and in exchange asks you to care about the five settings above.

For an internal app with bursty, business-hours traffic that can scale to zero at night, Cloud Run usually wins on both cost and attention. For something that has to keep state on local disk, run a long process that is not shaped like a request, or sit behind a fixed IP address for a vendor allowlist, a VM or a managed service is the cleaner answer, and we say so.

What we watch in the first month

  1. A billing budget with alerts on the project, set before the first deploy rather than after the first surprise.
  2. Instance count over time. A service that never returns to zero overnight is telling you something is polling it.
  3. Requests nobody ordered: anonymous hits, health checks from an unknown source, a client retrying in a loop.
  4. Cold start time, measured rather than assumed, so the decision about minimum instances rests on a number.
  5. Error rate against the new, lower timeout, every day for the first week after you change it.

How Guanacos Tech helps

We run this as a fixed piece of work: read the service as it stands today, set the billing mode, concurrency, instance cap and timeout to match the real user count, move secrets into Secret Manager, close the invoker grant, put a budget and alerts on the project, and leave a one-page note saying what each setting is and why it is where it is. We are an independent consultancy and our engineers hold Google certifications, so you get the reasoning along with the settings instead of a changed service and a shrug.

If you have an internal app on Cloud Run that nobody has opened since the day it shipped, that is the normal starting point and not an embarrassing one. See what we do in Google Cloud consulting for small teams. A 30 minute call is enough to say which of these settings is costing you money.

Sources

Next step

Would you rather we did this for you?

Thirty minutes on Google Meet, free. We look at your domain or project with you, tell you what is wrong and what we would do first. If you can fix it yourself, we say so.

Book a 30-minute call or read about our Google Cloud consulting

Frequently asked questions

Does Cloud Run cost anything while nobody is using the app?

Under the default request-based billing, idle instances that are not minimum instances are not charged, and an instance will not stay idle for more than 15 minutes after a request unless minimum instances keep it alive. If you set minimum instances above zero, those instances are billed at a reduced idle rate around the clock, including overnight.

How do I stop a Cloud Run service from being public?

Public access comes from granting the Cloud Run Invoker role to the allUsers principal. Remove that grant, give roles/run.invoker to the specific users or groups who need it, and deploy with the no-allow-unauthenticated flag so unauthenticated requests are rejected with an HTTP 403. Find out what is calling the service anonymously before you close it, because those callers will stop working immediately.

What should maximum instances be for a small internal app?

The default is 100 instances per revision, far more than a team of ten or twenty needs. A single-digit cap matched to the real user count turns a runaway retry loop into a visible error instead of a large bill, and you can raise it deliberately if real traffic ever calls for it.