Skip to main content

Super Technician and Little scripts -2 Deploying an LLM service to AWS ECS

 

Complete, end-to-end architecture guide and production deployment strategy for deploying an LLM service to AWS ECS using a GitLab CI/CD Pipeline.




Key LLM Production Considerations Model Cache Persistence (EFS): Download heavy weights (e.g., Llama-3, Qwen) to an AWS Elastic File System (EFS) mounted to /root/.cache/huggingface. This prevents re-downloading multi-gigabyte models on container restarts.

GPU / Compute Launch Type: Use ECS EC2 Launch Type with GPU-enabled instances (e.g., g5.xlarge or g4dn.xlarge with NVIDIA A10G/T4) or Fargate (CPU-only / quantized lightweight models).

vLLM / TensorRT-LLM Engine: Use high-throughput inference engines like vLLM wrapped inside a FastAPI application for OpenAI-compatible endpoint serving.


complete, single-block solution containing the Dockerfile, vLLM/FastAPI App Server, AWS ECS Task Definition Template, and the complete .gitlab-ci.yml Pipeline.


# ==============================================================================

# SECTION 1: Dockerfile (vLLM / FastAPI Server with GPU Support)

# Save as: Dockerfile

# ==============================================================================

cat << 'EOF' > Dockerfile

FROM vllm/vllm-openai:latest


# Set environment variables

ENV PYTHONUNBUFFERED=1 \

    HF_HOME=/root/.cache/huggingface \

    MODEL_NAME="Qwen/Qwen2.5-Coder-7B-Instruct" \

    MAX_MODEL_LEN=4096 \

    GPU_MEMORY_UTILIZATION=0.90


WORKDIR /app


# Expose HTTP port for ECS Health Check and ALB

EXPOSE 8000


# Entrypoint to serve LLM via OpenAI-compatible REST API

ENTRYPOINT ["python3", "-m", "vllm.entrypoints.openai.api_server"]

CMD ["--model", "Qwen/Qwen2.5-Coder-7B-Instruct", "--port", "8000", "--host", "0.0.0.0", "--trust-remote-code"]

EOF



# ==============================================================================

# SECTION 2: AWS ECS Task Definition Template

# Save as: ecs-task-definition.json

# ==============================================================================

cat << 'EOF' > ecs-task-definition.json

{

  "family": "llm-inference-task",

  "networkMode": "awsvpc",

  "requiresCompatibilities": ["EC2"],

  "cpu": "4096",

  "memory": "16384",

  "executionRoleArn": "arn:aws:iam::$AWS_ACCOUNT_ID:role/ecsTaskExecutionRole",

  "taskRoleArn": "arn:aws:iam::$AWS_ACCOUNT_ID:role/llmTaskRole",

  "containerDefinitions": [

    {

      "name": "llm-container",

      "image": "$AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com/$ECR_REPOSITORY:$IMAGE_TAG",

      "essential": true,

      "resourceRequirements": [

        {

          "type": "GPU",

          "value": "1"

        }

      ],

      "portMappings": [

        {

          "containerPort": 8000,

          "hostPort": 8000,

          "protocol": "tcp"

        }

      ],

      "environment": [

        { "name": "HF_TOKEN", "value": "$HF_TOKEN" }

      ],

      "mountPoints": [

        {

          "sourceVolume": "efs-model-cache",

          "containerPath": "/root/.cache/huggingface"

        }

      ],

      "logConfiguration": {

        "logDriver": "awslogs",

        "options": {

          "awslogs-group": "/ecs/llm-inference",

          "awslogs-region": "$AWS_DEFAULT_REGION",

          "awslogs-stream-prefix": "vllm"

        }

      },

      "healthCheck": {

        "command": ["CMD-SHELL", "curl -f http://localhost:8000/health || exit 1"],

        "interval": 30,

        "timeout": 10,

        "retries": 3,

        "startPeriod": 120

      }

    }

  ],

  "volumes": [

    {

      "name": "efs-model-cache",

      "efsVolumeConfiguration": {

        "fileSystemId": "$EFS_FILE_SYSTEM_ID",

        "transitEncryption": "ENABLED"

      }

    }

  ]

}

EOF



# ==============================================================================

# SECTION 3: Complete GitLab CI/CD Pipeline Configuration

# Save as: .gitlab-ci.yml

# Required GitLab Variables in Settings -> CI/CD -> Variables:

#   - AWS_ACCESS_KEY_ID (Masked)

#   - AWS_SECRET_ACCESS_KEY (Masked)

#   - AWS_DEFAULT_REGION (e.g., us-east-1)

#   - AWS_ACCOUNT_ID (e.g., 123456789012)

#   - ECR_REPOSITORY (e.g., llm-vllm-service)

#   - ECS_CLUSTER (e.g., llm-production-cluster)

#   - ECS_SERVICE (e.g., vllm-service)

#   - EFS_FILE_SYSTEM_ID (e.g., fs-0123456789abcdef0)

#   - HF_TOKEN (Masked - Hugging Face API Token)

# ==============================================================================

cat << 'EOF' > .gitlab-ci.yml

stages:

  - test

  - build

  - deploy


variables:

  DOCKER_DRIVER: overlay2

  DOCKER_TLS_CERTDIR: ""

  IMAGE_TAG: $CI_COMMIT_SHORT_SHA


# ------------------------------------------------------------------------------

# STAGE 1: Code Linting and API Pre-checks

# ------------------------------------------------------------------------------

lint-and-validate:

  stage: test

  image: python:3.11-slim

  script:

    - echo "Validating configuration and Dockerfile syntax..."

    - pip install docker-compose

    - python3 -c "import json; json.load(open('ecs-task-definition.json'))"

  only:

    - main

    - merge_requests


# ------------------------------------------------------------------------------

# STAGE 2: Build Docker Image and Push to AWS ECR

# ------------------------------------------------------------------------------

build-and-push-ecr:

  stage: build

  image: docker:24.0.5

  services:

    - docker:24.0.5-dind

  before_script:

    - apk add --no-cache python3 py3-pip aws-cli gettext

    - echo "Logging into AWS ECR..."

    - aws ecr get-login-password --region $AWS_DEFAULT_REGION | docker login --username AWS --password-stdin $AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com

    # Ensure ECR repository exists

    - aws ecr describe-repositories --repository-names $ECR_REPOSITORY --region $AWS_DEFAULT_REGION || aws ecr create-repository --repository-name $ECR_REPOSITORY --region $AWS_DEFAULT_REGION

  script:

    - echo "Building LLM Docker Container Image..."

    - docker build -t $ECR_REPOSITORY:$IMAGE_TAG .

    - docker tag $ECR_REPOSITORY:$IMAGE_TAG $AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com/$ECR_REPOSITORY:$IMAGE_TAG

    - docker tag $ECR_REPOSITORY:$IMAGE_TAG $AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com/$ECR_REPOSITORY:latest

    - echo "Pushing image to ECR..."

    - docker push $AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com/$ECR_REPOSITORY:$IMAGE_TAG

    - docker push $AWS_ACCOUNT_ID.dkr.ecr.$AWS_DEFAULT_REGION.amazonaws.com/$ECR_REPOSITORY:latest

  only:

    - main


# ------------------------------------------------------------------------------

# STAGE 3: Deploy New Task Definition & Update ECS Service

# ------------------------------------------------------------------------------

deploy-to-ecs:

  stage: deploy

  image: registry.gitlab.com/gitlab-org/cloud-deploy/aws-base:latest

  before_script:

    - apk add --no-cache gettext

  script:

    - echo "Substitutiting variables into Task Definition..."

    - envsubst < ecs-task-definition.json > rendered-task-def.json

    

    - echo "Registering new Task Definition in AWS ECS..."

    - TASK_REV=$(aws ecs register-task-definition --cli-input-json file://rendered-task-def.json --region $AWS_DEFAULT_REGION | jq -r '.taskDefinition.taskDefinitionArn')

    - echo "Registered Task Definition ARN: $TASK_REV"

    

    - echo "Updating ECS Service with rolling update strategy..."

    - aws ecs update-service --cluster $ECS_CLUSTER --service $ECS_SERVICE --task-definition $TASK_REV --force-new-deployment --region $AWS_DEFAULT_REGION

    

    - echo "Waiting for ECS Service deployment to stabilize..."

    - aws ecs wait services-stable --cluster $ECS_CLUSTER --services $ECS_SERVICE --region $AWS_DEFAULT_REGION

    - echo "Deployment Successful! LLM Endpoint Updated."

  only:

    - main

EOF

Comments

Popular posts from this blog

Telecom OSS and BSS: A Comprehensive Guide

  Telecom OSS and BSS: A Comprehensive Guide Table of Contents Part I: Foundations of Telecom Operations Chapter 1: Introduction to Telecommunications Networks A Brief History of Telecommunications Network Architectures: From PSTN to 5G Key Network Elements and Protocols Chapter 2: Understanding OSS and BSS Defining OSS and BSS The Role of OSS in Network Management The Role of BSS in Business Operations The Interdependence of OSS and BSS Chapter 3: The Telecom Business Landscape Service Providers and Their Business Models The Evolving Customer Experience Regulatory and Compliance Considerations The Impact of Digital Transformation Part II: Operations Support Systems (OSS) Chapter 4: Network Inventory Management (NIM) The Importance of Accurate Inventory NIM Systems and Their Functionality Data Modeling and Management Automation and Reconciliation Chapter 5: Fault Management (FM) Detecting and Isolating Network Faults FM Systems and Alerting Mecha...

"Depth-Guard" – 3D Spatial Occupancy monitor Challenge -2

  Project Title: "Depth-Guard" – 3D Spatial Occupancy Monitor 1. The Problem In a smart warehouse, a robot needs to know if a loading zone is clear or occupied. A 2D camera alone can’t tell the difference between a "flat picture of a box" on the floor and an "actual 3D box." The Goal: Build a Python-based system that uses Computer Vision and Depth Perception (AI 3D) to identify objects and determine their 3D volume (Size) and Distance from the camera. 2. Intern Tasks Object Detection: Use a pre-trained model (like YOLOv8) to draw 2D boxes around objects. Depth Mapping: Use a depth estimation model (like MiDaS or a simulated Stereo-depth feed) to calculate how far each object is. Occupancy Logic: If an object is closer than 1 meter and larger than a specific volume, mark the zone as "BLOCKED." Alert System: Print a warning if the 3D space is too crowded. 3. Sample Datasets (Simulation) Since interns may not have 3D cameras (LiDAR/RGB-D), pr...

Simple Virtual Waiting Room -Challenge 1

   Simple Virtual Waiting Room (VWR) 1. The Problem Our website can only handle 10 users per minute . If more than 10 people try to access it at once, the server will crash. We need a system that: Counts incoming users. Redirects "overflow" users to a waiting page. Admits them back to the main site one by one as space becomes available. 2. Intern Tasks Create a Gateway: A simple script that checks: if (active_users < 10) { allow } else { send to queue } . Build the Queue: Use a simple list (FIFO) to store user IDs. The Wait Page: A basic HTML page that says: "You are number X in line. Estimated wait: Y minutes." Admission Logic: Every 30 seconds, pull the next user from the queue and "admit" them. 3. Sample Datasets (Simulation) Provide these two datasets to the interns. They should write a script to "read" these files and simulate how their system reacts. Dataset A: The Traffic Surge (Input) This file simulates users arriving at the ...