Skip to main content

Posts

AI Operating Systems (AIOS)

  AI Operating Systems (AIOS) & NetObserve: Unifying Agentic Swarms with Ring-0 Telemetry Part I: The AIOS Revolution & Trillion-Dollar Paradigm Chapter 1: The Paradigm Shift — From Stateless Monolithic LLM APIs to Stateful AI Operating Systems Chapter 2: Anatomy of an AIOS Kernel — Agent Schedulers, KV-Cache Management & OS Context Switching Chapter 3: The Trillion-Dollar Value Drivers — Multi-Agent Swarms & Enterprise Automation Frameworks Part II: NetObserve Fundamentals & Ring-0 Network Telemetry Chapter 4: Zero-Overhead Kernel Observability — Hooking eBPF into sys_enter_connect & Linux Socket Buffers Chapter 5: Microsecond Network Telemetry — Passive TCP RTT, Retransmission Tracking & EC2 Hypervisor Limits Chapter 6: Container Mesh Visibility — Mapping EKS East-West Pod Topologies Without Sidecar Bloat Part III: Unifying L7 Agent Swarms with Ring-0 Observability Chapter 7: End-to-End Context Tracing — Correlating Agent Swarm IDs & Prompt...
Recent posts

Super Technician and Little scripts -2 Deploying an LLM service to AWS ECS

  Complete, end-to-end architecture guide and production deployment strategy for deploying an LLM service to AWS ECS using a GitLab CI/CD Pipeline. Key LLM Production Considerations Model Cache Persistence (EFS): Download heavy weights (e.g., Llama-3, Qwen) to an AWS Elastic File System (EFS) mounted to /root/.cache/huggingface. This prevents re-downloading multi-gigabyte models on container restarts. GPU / Compute Launch Type: Use ECS EC2 Launch Type with GPU-enabled instances (e.g., g5.xlarge or g4dn.xlarge with NVIDIA A10G/T4) or Fargate (CPU-only / quantized lightweight models). vLLM / TensorRT-LLM Engine: Use high-throughput inference engines like vLLM wrapped inside a FastAPI application for OpenAI-compatible endpoint serving. complete, single-block solution containing the Dockerfile, vLLM/FastAPI App Server, AWS ECS Task Definition Template, and the complete .gitlab-ci.yml Pipeline. # ============================================================================== # SECTION ...

Super Technician and little scripts : Server Health & Storage Management

  Server Health & Storage Management: Automated Cleanup, Disk Offloading, and MySQL Backups When running high-traffic web applications, database services, and build environments, server disks can rapidly fill up due to log aggregation, build caches, and upload directories. Reaching 100% disk usage on the root partition (/) often results in service crashes, database corruption, and deployment failures. This guide details a multi-tiered strategy for managing server storage using secondary disks, directory symlinking, log truncation, and scheduled database backups. Key Architecture Strategies Offloading Storage via Symlinks Instead of reconfiguring internal application paths, heavy folders (such as file upload directories, media storage, or virtual environments) can be moved to an external or secondary mount point (e.g., /data). By creating a symbolic link (ln -s) from the original directory path pointing to the secondary disk, applications continue reading and writing seamlessl...