GitHub Actions 容器作业与 Service 容器:job container、健康检查与集成测试

GitHub Actions 容器作业与 Service 容器实战:jobs.container 语法(image、credentials、env、ports、volumes、options)、services 定义 PostgreSQL 与 Redis、健康检查与就绪等待、网络模型与 localhost 踩坑、集成测试实战、容器内 checkout 与缓存、私有镜像认证与性能优化

单元测试跑得再快,也拦不住「连不上数据库」这类集成问题。GitHub Actions 的容器作业(job container)与 Service 容器(service container)让 CI 能在完全受控的镜像里运行测试,并顺手拉起一套 PostgreSQL、Redis 供集成测试使用。但端口映射、网络别名、健康检查与就绪等待这些细节,恰恰是踩坑最密集的地方。本文从语法到实战,把容器作业与 Service 容器的每一个参数讲清楚。


一、容器作业基础:jobs.container 语法

1.1 为什么要把 job 放进容器

默认情况下,job 直接跑在 runner 的虚拟机上,工具链版本取决于 runner 镜像预装了什么。容器作业则把整个 job 放进一个你指定的镜像里执行,带来三点好处:

可复现:镜像 tag 固定,工具链版本不随 runner 镜像升级而漂移
隔离:依赖装在容器内,不污染 runner,也不需要 sudo 装包
一致:本地 docker run 能复现的,CI 里也能复现

代价是每次都要拉取镜像,且容器内默认缺少一些 runner 预装工具(如 git、curl),需要自己补。

1.2 container 语法全貌

# .github/workflows/container-job.yml
name: Container Job
on: [push, pull_request]

jobs:
  test:
    runs-on: ubuntu-latest
    container:
      image: node:20-bookworm
      credentials:
        username: ${{ github.actor }}
        password: ${{ secrets.GITHUB_TOKEN }}
      env:
        NODE_ENV: test
        DATABASE_URL: postgres://postgres:postgres@postgres:5432/test
      ports:
        - 8080:80
      volumes:
        - /tmp/cache:/cache
      options: --cpus 2 --memory 4g --health-cmd "node --version"
    steps:
      - uses: actions/checkout@v4
      - run: node --version

逐字段说明:

image       必填,镜像名(可带 tag/digest),如 ghcr.io/owner/img:v1
credentials 私有仓库认证,username/password
env         注入到容器内的环境变量,覆盖镜像自带的同名变量
ports       把容器端口映射到宿主机,格式 宿主机:容器
volumes   挂载卷,常用于缓存与工作目录
options   直接透传给 docker create 的参数,是能力最强也最容易写错的字段

1.3 options 常用参数

options 是原样拼进 docker create 命令的字符串,因此 Docker 的常用开关都能用:

    container:
      image: ubuntu:22.04
      options: >-
        --cpus 2
        --memory 4g
        --health-cmd "pg_isready -U postgres"
        --health-interval 10s
        --health-timeout 5s
        --health-retries 5

注意 options 里的引号会被 shell 解析,含空格的值务必用引号包住,否则会被拆成多个参数导致 docker create 报错。


二、Service 容器:数据库与中间件

2.1 services 语法

services 声明的容器与 job 并行启动,job 结束后自动销毁,适合放测试依赖的数据库与中间件。

jobs:
  test:
    runs-on: ubuntu-latest
    services:
      postgres:
        image: postgres:16
        env:
          POSTGRES_USER: postgres
          POSTGRES_PASSWORD: postgres
          POSTGRES_DB: test
        ports:
          - 5432:5432
        options: >-
          --health-cmd pg_isready
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5

service 与 container 共享同一套子字段:image、credentials、env、ports、volumes、options。区别在于 services 可以写多个,而 job 的 container 只能有一个。

2.2 PostgreSQL 与 MySQL

    services:
      postgres:
        image: postgres:16
        env:
          POSTGRES_PASSWORD: postgres
        ports:
          - 5432:5432
        options: >-
          --health-cmd pg_isready
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5
      mysql:
        image: mysql:8.0
        env:
          MYSQL_ROOT_PASSWORD: root
          MYSQL_DATABASE: test
        ports:
          - 3306:3306
        options: >-
          --health-cmd "mysqladmin ping -h 127.0.0.1"
          --health-interval 10s
          --health-timeout 5s
          --health-retries 10

MySQL 首次启动初始化数据目录较慢,--health-retries 要给足,否则健康检查未过就进入步骤,连接必然失败。

2.3 Redis 与 RabbitMQ

    services:
      redis:
        image: redis:7
        ports:
          - 6379:6379
        options: >-
          --health-cmd "redis-cli ping"
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5
      rabbitmq:
        image: rabbitmq:3.13-management
        env:
          RABBITMQ_DEFAULT_USER: guest
          RABBITMQ_DEFAULT_PASS: guest
        ports:
          - 5672:5672
          - 15672:15672
        options: >-
          --health-cmd "rabbitmq-diagnostics -q ping"
          --health-interval 10s
          --health-timeout 5s
          --health-retries 10

RabbitMQ 的 rabbitmq-diagnostics ping 比 rabbitmqctl status 更轻量,适合做健康探针。

2.4 Elasticsearch

      elasticsearch:
        image: elasticsearch:8.13.0
        env:
          discovery.type: single-node
          xpack.security.enabled: "false"
          ES_JAVA_OPTS: "-Xms512m -Xmx512m"
        ports:
          - 9200:9200
        options: >-
          --health-cmd "curl -sf http://localhost:9200/_cluster/health"
          --health-interval 15s
          --health-timeout 10s
          --health-retries 10

ES 内存占用大,务必用 ES_JAVA_OPTS 限制堆大小,否则在 7G 内存的 runner 上容易触发 OOM 被内核杀掉。


三、网络与就绪等待

3.1 网络模型:localhost 还是服务名

这是容器作业最经典的坑。答案取决于 job 本身是否运行在容器里:

job 直接跑在 runner 上(无 container 字段)
  → 通过 localhost:端口 访问 service,因为端口被映射到了宿主机

job 运行在容器里(有 container 字段)
  → service 与 job 同在 user-defined bridge 网络
  → 用「服务名:容器端口」访问,如 postgres:5432
  → 此时 localhost 指向 job 容器自身,连不上 service

也就是说,同一份 services 定义,在加不加 container 时连接串完全不同:

    container:
      image: node:20
    services:
      postgres:
        image: postgres:16
        ports:
          - 5432:5432
    steps:
      - run: |
          # 容器作业内:必须用服务名
          psql postgres://postgres:postgres@postgres:5432/test -c "select 1"

3.2 健康检查 options

健康检查是「声明式等待」:Docker 会周期性执行 --health-cmd,直到状态变为 healthy,Actions 才启动 job 步骤。

--health-cmd      探针命令,返回 0 视为健康
--health-interval 两次探针间隔,默认 30s
--health-timeout  单次探针超时
--health-retries  连续失败多少次后标记 unhealthy
--health-start-period 启动宽限期,此期间失败不计入 retries

对启动慢的中间件,--health-start-period 比单纯加大 --health-retries 更精确,因为它把「还在初始化」与「真的坏了」区分开。

3.3 wait-for-it 与重试循环

健康检查并非万能:某些镜像本身不带探针工具,或服务在 healthy 之后仍需数秒完成初始化。兜底方案是显式等待。

      - name: Wait for Postgres
        run: |
          for i in $(seq 1 30); do
            if pg_isready -h postgres -p 5432 -U postgres; then
              echo "postgres is ready"
              exit 0
            fi
            echo "waiting for postgres ($i/30)"
            sleep 2
          done
          echo "postgres did not become ready in time" >&2
          exit 1

另一种做法是把 wait-for-it.sh 或 dockerize 作为脚本引入容器,用 ./wait-for-it.sh postgres:5432 -- npm test 包住测试命令。两者本质相同,都是「带超时的轮询」。


四、集成测试实战

4.1 Laravel 连 MySQL 跑测试

name: Laravel Tests
on: [push, pull_request]

jobs:
  test:
    runs-on: ubuntu-latest
    container:
      image: php:8.3-cli
      options: --health-cmd "php -v"
    services:
      mysql:
        image: mysql:8.0
        env:
          MYSQL_ROOT_PASSWORD: root
          MYSQL_DATABASE: testing
        ports:
          - 3306:3306
        options: >-
          --health-cmd "mysqladmin ping -h 127.0.0.1"
          --health-interval 10s
          --health-retries 10
    env:
      DB_CONNECTION: mysql
      DB_HOST: mysql
      DB_PORT: 3306
      DB_DATABASE: testing
      DB_USERNAME: root
      DB_PASSWORD: root
    steps:
      - uses: actions/checkout@v4
      - name: Install extensions
        run: |
          docker-php-ext-install pdo_mysql
      - name: Install dependencies
        run: |
          curl -sS https://getcomposer.org/installer | php -- --install-dir=/usr/local/bin --filename=composer
          composer install --prefer-dist --no-interaction
      - name: Run tests
        run: php artisan test

注意 DB_HOST: mysql 用的是服务名,这正是 3.1 节结论的落地。

4.2 Spring Boot 连 PostgreSQL

    services:
      postgres:
        image: postgres:16
        env:
          POSTGRES_PASSWORD: postgres
        ports:
          - 5432:5432
        options: >-
          --health-cmd pg_isready
          --health-interval 10s
          --health-retries 5
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-java@v4
        with:
          distribution: temurin
          java-version: "21"
          cache: maven
      - name: Run integration tests
        run: mvn -B verify
        env:
          SPRING_DATASOURCE_URL: jdbc:postgresql://localhost:5432/postgres
          SPRING_DATASOURCE_USERNAME: postgres
          SPRING_DATASOURCE_PASSWORD: postgres

这里 job 没有 container 字段,所以走 localhost:5432。同一篇文章里两种写法并存,正是为了强调差异。

4.3 Node.js 连 Redis 跑接口测试

    services:
      redis:
        image: redis:7
        ports:
          - 6379:6379
        options: >-
          --health-cmd "redis-cli ping"
          --health-interval 10s
          --health-retries 5
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: "20"
          cache: npm
      - run: npm ci
      - run: npm run test:integration
        env:
          REDIS_URL: redis://localhost:6379

五、容器内的缓存与 checkout

5.1 actions/checkout 与 /github/workspace

容器作业里,runner 会把工作目录挂载到容器的 /github/workspace,actions/checkout 也把代码放到这里,所以 run 步骤的默认工作目录就是它,无需手动 cd。

/github/workspace  代码与默认工作目录
/github/home       容器内的 HOME,缓存目录常放这里
/github/workflow   工作流事件载荷

5.2 容器内缺少 git 的坑

actions/checkout 依赖容器内存在 git。官方精简镜像(如 php:8.3-cli、node:20-alpine)往往不带 git,checkout 会直接失败。解决办法有两种:

      # 方案一:先装 git,再 checkout
      - name: Install git
        run: apt-get update && apt-get install -y git
      - uses: actions/checkout@v4

      # 方案二:用自带 git 的基础镜像
      # container:
      #   image: node:20-bookworm   # bookworm 完整版带 git

方案二更省时间,但镜像体积更大。若只是为了跑测试,推荐方案一配合 slim 镜像。

5.3 缓存挂载

actions/cache 在容器作业里同样可用,但缓存路径要在容器内可见。用 volumes 把宿主机目录挂进容器做持久缓存是另一种思路:

    container:
      image: node:20-bookworm
      volumes:
        - /home/runner/.npm:/root/.npm

注意 volumes 用的是宿主机绝对路径,容器作业的宿主机就是 runner 虚拟机。这条路径在 job 结束后不保留,因此只适合单 job 内的复用,跨 job 复用仍应交给 actions/cache。


六、私有镜像认证与性能可靠性

6.1 credentials 与 GHCR

拉取私有镜像需要认证。GitHub Container Registry(GHCR)可以直接用内置的 GITHUB_TOKEN,无需额外配置 secret:

    container:
      image: ghcr.io/my-org/private-image:latest
      credentials:
        username: ${{ github.actor }}
        password: ${{ secrets.GITHUB_TOKEN }}

前提是工作流已声明 packages: read 权限:

permissions:
  contents: read
  packages: read

对于 Docker Hub 或自建 registry,把 username 换成 secrets.DOCKERHUB_USERNAME、password 换成对应的 token 即可。services 里的私有镜像同样支持 credentials 字段。

6.2 镜像拉取优化

容器作业最常见的性能问题是镜像拉取慢。可操作的手段:

固定 digest 而非 latest:latest 会导致缓存命中不稳定
精简基础镜像:slim / alpine 拉取更快,但要补装 git 等工具
自建 runner 预热:self-hosted runner 上镜像常驻,拉取近乎为零
合并 job:一次拉取跑完所有步骤,胜过多个 job 各拉一次

6.3 资源限制与 matrix 版本

用 options 给容器设上限,避免个别 job 吃满 runner 内存拖垮整机:

    strategy:
      matrix:
        postgres: ["14", "15", "16"]
    services:
      postgres:
        image: postgres:${{ matrix.postgres }}
        env:
          POSTGRES_PASSWORD: postgres
        ports:
          - 5432:5432
        options: >-
          --health-cmd pg_isready
          --health-retries 5

矩阵变量在 services 中同样可插值,这让「多数据库版本兼容性测试」只需要一份配置。注意 ports 在所有矩阵分支里都是 5432:5432,因为每个分支跑在独立的 runner 上,不会冲突。


七、速查表

【连接地址】
job 无 container → localhost:映射端口
job 有 container → 服务名:容器端口

【container 字段】
image / credentials / env / ports / volumes / options

【services 字段】
与 container 相同,但可定义多个

【健康检查 options】
--health-cmd / --health-interval / --health-timeout
--health-retries / --health-start-period

【探针命令参考】
postgres     pg_isready
mysql        mysqladmin ping -h 127.0.0.1
redis        redis-cli ping
rabbitmq     rabbitmq-diagnostics -q ping
elasticsearch curl -sf http://localhost:9200/_cluster/health

【路径】
代码目录     /github/workspace
HOME         /github/home

【私有镜像】
credentials.username / credentials.password
GHCR 用 github.actor + secrets.GITHUB_TOKEN

总结

容器作业与 Service 容器解决的是同一个问题的两面:前者让 job 跑在可复现的镜像里,后者让依赖服务随 job 按需拉起。掌握三个关键点就能少踩绝大多数坑——第一,连接地址取决于 job 是否在容器内,无 container 用 localhost,有 container 用服务名;第二,用 --health-cmd 加 --health-retries 做声明式等待,慢启动的中间件再叠加 --health-start-period 或轮询兜底;第三,容器内 actions/checkout 依赖 git,精简镜像要先补装。性能上优先固定镜像 digest、精简基础镜像,有条件用 self-hosted runner 预热。把这些配置固化进模板,集成测试就能从「本地能过 CI 不过」变成稳定可信的一道关卡。

延伸阅读:

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「github-actions」更多文章

  1. GitHub Actions 本地调试与排错:act、workflow_dispatch、日志与重跑
  2. GitHub Actions 工作流性能与并发控制:concurrency、超时与分钟数优化
  3. GitHub Actions 分支保护与 Rulesets:必需检查、CODEOWNERS 与仓库治理