# Amir Teymoori > AI and LLM Engineer specializing in Android, Python, and full-stack development. Experienced in building intelligent systems, automation tools, and scalable backend APIs. Passionate about applied machine learning, drone technology, and creative software design. ## About This Site This is Amir Teymoori's technical blog focused on AI, LLMs, machine learning, and software development. Articles cover practical implementations, tutorials, comparisons, and best practices for AI engineers and developers. ## Author - **Name:** Amir Teymoori - **Role:** AI & LLM Engineer - **Expertise:** Large Language Models, RAG Systems, AI Agents, Python, TypeScript, Android/Kotlin - **Website:** https://amirteymoori.com/ - **GitHub:** https://github.com/ateymoori/ - **LinkedIn:** https://www.linkedin.com/in/amirhossein-teymoori/ - **Twitter/X:** https://x.com/AmirHosseiin ## Topics Covered - AI Coding Tools: Claude Code, Cursor, Windsurf, OpenCode - LLM Development: Prompt Engineering, Fine-tuning, Evaluation - RAG & Vector Databases: Embeddings, Chunking Strategies, Retrieval - AI Agents & MCP Servers: Tool Use, Multi-agent Systems - MLOps & Observability: LangFuse, Evaluation Pipelines - DevOps for AI: Docker, Deployment, Cost Optimization ## Sitemaps & Feeds - **XML Sitemap:** https://amirteymoori.com/sitemap_index.xml - **RSS Feed:** https://amirteymoori.com/feed/ - **Markdown Posts:** https://amirteymoori.com/posts/ (individual .md files) ## Categories ### LLM Models, Providers and Training Large language models in practice. GPT-4.x/ChatGPT (OpenAI), Claude, Gemini, Grok, Llama, Mixtral, Qwen, Flux, Cerebras-GPT, Hugging Face models. Tokenization, embeddings, KV cache, context windows. Fine-tuning with LoRA/QLoRA and PEFT. Persian-focused runs and data cleaning. Compare APIs, pricing, and rate limits. Reproducible notebooks and CLI. - URL: https://amirteymoori.com/category/llm-models-providers-training/ - Posts: 13 ### Vibe-Coding: Cursor, Claude Code, Windsurf Ship features with AI IDEs. Cursor, Claude Code, Windsurf, VS Code, GitHub Copilot, OpenCode, Kilo Code. Context windows, repo maps, test scaffolds, and refactor loops. "Prompt to PR" flow, review checklists, and prompt libraries. Live coding on macOS with Warp. Branch strategy, snapshots, and integration with LangSmith. - URL: https://amirteymoori.com/category/vibe-coding-cursor-claude-code-windsurf/ - Posts: 7 ### Prompt and Context Engineering Patterns for prompt engineering and context engineering that convert. Few-shot, system prompts, tool prompts, and safety. Context packing, chunking, reranking, and window budgets. Hands-on with ChatGPT (OpenAI), Claude Code, Gemini, Grok, and Groq API. Orchestrate with LangChain, LangGraph, DSPy, and LangSmith. Route via Portkey or OpenRouter. Code samples for JSON tools, retries, and eval loops. - URL: https://amirteymoori.com/category/prompt-context-engineering/ - Posts: 6 ### AI Agents, Tools and MCP Servers Plan, act, and use tools. ReAct, function calling, memory, and auditing. Build MCP servers (FastMCP) for Chrome DevTools, Spotify, and custom REST. Define JSON schemas, auth, rate limits, and retries. Integrate with Cursor, Windsurf, VS Code, and Claude Code. Ship assistants that call search, code, and data tools safely. - URL: https://amirteymoori.com/category/ai-agents-tools-mcp-servers/ - Posts: 5 ### RAG, Graph RAG and Vector Databases Build RAG that answers. Hybrid search, rerankers, caching, and freshness. Graph RAG with entity and relation extraction. Vector DBs: Typesense, Qdrant, FAISS, Milvus, pgvector, Supabase Vector. Chunking recipes, metadata, and dedupe. Wire to LangChain, DSPy, and LangGraph. Benchmarks and latency budgets with real datasets. - URL: https://amirteymoori.com/category/rag-graph-rag-vector-databases/ - Posts: 3 ### Inference, Serving and Cost Control Serve models fast and cheap. vLLM, TGI, Ollama, Groq LPU, OpenAI API, OpenRouter, Portkey gateways. Streaming, batching, KV cache pinning, quantization (GGUF, AWQ). Autoscale, canary, and circuit breakers. FastAPI gateways, gRPC, and WebSockets. Mac OS local dev with Warp and Docker. Helm charts and Terraform snippets. - URL: https://amirteymoori.com/category/inference-serving-cost-control/ - Posts: 3 ### MLOps, Evaluation and Observability Track quality and regressions. MLflow for runs and params. LangSmith for traces and evals. Custom evaluators in DSPy. Metrics: accuracy, latency, cost, toxicity. Dashboards with Prometheus, Grafana, and Sentry. Data versioning, synthetic data, and red-team sets. Ship/rollback rules with thresholds. - URL: https://amirteymoori.com/category/mlops-evaluation-observability/ - Posts: 3 ### DevOps for AI: Docker, Caddy, CI/CD Reproducible deploys for LLM apps. Multi-stage Dockerfiles, health checks, and secrets. Caddy reverse proxy with TLS and HTTP/2. GitHub Actions, runners, and blue-green. Supabase, Postgres, Redis, and object storage. Ubuntu hardening, Cloudflare, and logs. From Mac Studio to VPS with one-command releases. - URL: https://amirteymoori.com/category/devops-ai-docker-caddy-ci-cd/ - Posts: 2 ### Benchmarks, Case Studies and Playbooks Real builds end to end. RAG vs Graph RAG comparisons. Provider shootouts: OpenAI, Anthropic, Google, Groq, xAI, Cerebras, Hugging Face. Load tests with k6 and wrk. Cost per 1k tokens, p95 latency, and throughput. Postmortems, failure modes, and checklists. Templates to copy for new projects. - URL: https://amirteymoori.com/category/benchmarks-case-studies-playbooks/ - Posts: 2 ### Open-Source Projects Open-source software projects and applications - URL: https://amirteymoori.com/category/open-source-projects/ - Posts: 1 ## Recent Articles ### LLM Guardrails That Survive Production in 2026 - **URL:** https://amirteymoori.com/llm-guardrails-prompt-injection-pii-agents-production/ - **Markdown:** https://amirteymoori.com/posts/llm-guardrails-prompt-injection-pii-agents-production.md - **Date:** 2026-08-03 - **Category:** MLOps, Evaluation and Observability - **Tags:** 2026, ai-engineering, ai-security, evals, llm-agents, llm-guardrails, owasp, pii, Prompt Injection - **Summary:** Prompt injection, PII leaks, and agents calling the wrong tool. The layered guardrail stack I actually run in production, plus evals that keep it honest. ### LLM API Pricing 2026: Every Major Model Compared - **URL:** https://amirteymoori.com/llm-api-pricing-2026-claude-gpt-deepseek-qwen-comparison/ - **Markdown:** https://amirteymoori.com/posts/llm-api-pricing-2026-claude-gpt-deepseek-qwen-comparison.md - **Date:** 2026-08-03 - **Category:** Inference, Serving and Cost Control - **Tags:** 2026, api-costs, claude, cost optimization, deepseek, gemini, GPT, LLM comparison, llm-pricing, qwen - **Summary:** Every major LLM API priced per million tokens as of August 3, 2026: Claude, GPT-5.6, Gemini, DeepSeek V4 and Qwen, plus the 3 patterns that cut bills. ### Small LLMs in 2026: When 7B Beats Last Year's 70B - **URL:** https://amirteymoori.com/small-llms-7b-on-device-qwen-gemma-efficiency-2026/ - **Markdown:** https://amirteymoori.com/posts/small-llms-7b-on-device-qwen-gemma-efficiency-2026.md - **Date:** 2026-08-03 - **Category:** Inference, Serving and Cost Control - **Tags:** 2026, gemma, llama-cpp, local-llm, model-routing, ollama, on-device, quantization, qwen, small-llms - **Summary:** A 7B model now does what 70B did last year. What small models handle in 2026, where they fail, and how to route around it. ### AI Agent Memory: What Actually Works in 2026 - **URL:** https://amirteymoori.com/ai-agent-memory-markdown-files-vs-vector-mem0-2026/ - **Markdown:** https://amirteymoori.com/posts/ai-agent-memory-markdown-files-vs-vector-mem0-2026.md - **Date:** 2026-08-03 - **Category:** AI Agents, Tools and MCP Servers - **Tags:** 2026, agent-memory, AI agents, context-engineering, letta, markdown, mem0, rag, vector-database, zep - **Summary:** Markdown files beat most agent memory products. Here's when file-based memory wins, and the exact scale where Mem0, Zep, or Letta earn their keep. ### DeepSeek V4: Open Weights Reach the Frontier - **URL:** https://amirteymoori.com/deepseek-v4-flash-open-weight-frontier-llm-review/ - **Markdown:** https://amirteymoori.com/posts/deepseek-v4-flash-open-weight-frontier-llm-review.md - **Date:** 2026-08-03 - **Category:** LLM Models, Providers and Training - **Tags:** 2026, ai-coding, deepseek, deepseek-v4, LLM comparison, moe, Open Source, open-weights, self-hosting - **Summary:** DeepSeek V4 ships MIT-licensed weights at 1.6T params and a 1M context. A developer review of the specs, the price, and what you can actually run. ### Sandbox AI Coding Agents Before They Sandbox You - **URL:** https://amirteymoori.com/sandbox-ai-coding-agents-docker-vm-e2b-security/ - **Markdown:** https://amirteymoori.com/posts/sandbox-ai-coding-agents-docker-vm-e2b-security.md - **Date:** 2026-08-03 - **Category:** DevOps for AI: Docker, Caddy, CI/CD - **Tags:** 2026, AI agents, Claude Code, devcontainers, Docker, e2b, firecracker, Prompt Injection, sandbox, security - **Summary:** Most developers disabled agent permission prompts on day three. Here is the real risk model and a sandboxing ladder, from allowlists to microVMs. ### 1M-Token Context: Do You Still Need RAG? - **URL:** https://amirteymoori.com/1m-token-context-window-vs-rag-llm-2026/ - **Markdown:** https://amirteymoori.com/posts/1m-token-context-window-vs-rag-llm-2026.md - **Date:** 2026-08-03 - **Category:** RAG, Graph RAG and Vector Databases - **Tags:** 1m-tokens, 2026, context window, context-engineering, hybrid-rag, llm-cost-optimization, long-context, rag, vector-search - **Summary:** RAG isn't dead, but your 2024 architecture is. Real cost math on 1M-token contexts, where recall breaks, and the hybrid pattern that won. ### OpenCode vs Claude Code: 2026 CLI Comparison - **URL:** https://amirteymoori.com/opencode-vs-claude-code-ai-cli-agent-comparison-2026/ - **Markdown:** https://amirteymoori.com/posts/opencode-vs-claude-code-ai-cli-agent-comparison-2026.md - **Date:** 2026-08-03 - **Category:** Vibe-Coding: Cursor, Claude Code, Windsurf - **Tags:** 2026, ai-coding, Claude Code, cli-tools, coding-agents, comparison, developer-tools, OpenCode - **Summary:** I run Claude Code and OpenCode side by side every day. Here is the honest 2026 comparison: real prices, real annoyances, and who should pick which. ### Claude Code Subagents: Multi-Agent Setup Guide - **URL:** https://amirteymoori.com/claude-code-subagents-skills-multi-agent-setup/ - **Markdown:** https://amirteymoori.com/posts/claude-code-subagents-skills-multi-agent-setup.md - **Date:** 2026-08-03 - **Category:** Vibe-Coding: Cursor, Claude Code, Windsurf - **Tags:** 2026, agentic-coding, AI agents, anthropic, Claude Code, claude-skills, developer-tools, multi-agent, subagents - **Summary:** My Claude Code setup: four subagent files, two skills, cheap models for search, and background tasks. Here is the exact config I run daily. ### Claude Fable 5 vs GPT-5.6: AI Coding Compared - **URL:** https://amirteymoori.com/claude-fable-5-mythos-vs-gpt-5-6-ai-coding-comparison/ - **Markdown:** https://amirteymoori.com/posts/claude-fable-5-mythos-vs-gpt-5-6-ai-coding-comparison.md - **Date:** 2026-08-03 - **Category:** LLM Models, Providers and Training - **Tags:** 2026, ai-coding, anthropic, claude-fable-5, claude-opus-5, gpt-5-6, LLM comparison, llm-pricing, mythos, swe-bench - **Summary:** Claude Fable 5 vs GPT-5.6 for real coding in August 2026: the export-ban saga, $10/$50 vs $5/$30 pricing, benchmarks, and which one I actually use. ### Fine-Tune Gemma 4 E4B for Swedish Translation - **URL:** https://amirteymoori.com/fine-tune-gemma-4-e4b-english-swedish-translation/ - **Markdown:** https://amirteymoori.com/posts/fine-tune-gemma-4-e4b-english-swedish-translation.md - **Date:** 2026-05-24 - **Category:** LLM Models, Providers and Training - **Tags:** 2026, Claude Opus, comet-score, fine-tuning, gemma, gemma-4, llm-as-judge, LoRA, machine-translation, swedish-translation, unsloth - **Summary:** Fine-tune Gemma 4 E4B with Unsloth and OPUS Books for English-Swedish novel translation. LoRA, BLEU, chrF, COMET, and Claude Opus 4.7 as judge. ### LLM-as-Judge Evals for Support AI Agents - **URL:** https://amirteymoori.com/llm-as-judge-eval-pipelines-customer-support-ai/ - **Markdown:** https://amirteymoori.com/posts/llm-as-judge-eval-pipelines-customer-support-ai.md - **Date:** 2026-05-23 - **Category:** MLOps, Evaluation and Observability - **Tags:** 2026, AI agents, customer-support-ai, evals, guardrails, llm-as-judge, MLOps, observability, prompt engineering - **Summary:** Build an LLM-as-judge eval pipeline for customer support AI agents using golden datasets, traces, rubrics, guardrails, and release gates. ### Hermes Agent vs OpenClaw: 2026 Comparison - **URL:** https://amirteymoori.com/hermes-agent-vs-openclaw-ai-assistant-comparison/ - **Markdown:** https://amirteymoori.com/posts/hermes-agent-vs-openclaw-ai-assistant-comparison.md - **Date:** 2026-05-01 - **Category:** AI Agents, Tools and MCP Servers - **Tags:** AI agents, AI Assistant, AI Comparison, Automation, Hermes Agent, MCP, Nous Research, Open Source, OpenClaw, Self-Improving AI - **Summary:** Compare Hermes Agent and OpenClaw, the two top open-source AI assistants of 2026. Architecture, security, costs, and migration paths covered. ### How LLM Chatbots Get Hacked: Prompt Injection - **URL:** https://amirteymoori.com/hack-llm-chatbot-extract-system-prompt-identify-ai-model/ - **Markdown:** https://amirteymoori.com/posts/hack-llm-chatbot-extract-system-prompt-identify-ai-model.md - **Date:** 2026-02-19 - **Category:** Prompt and Context Engineering - **Tags:** ai-safety, ai-security, chatbot-hacking, llm-fingerprinting, llm-security, owasp, Prompt Injection, prompt-extraction, red-teaming, system-prompt - **Summary:** A practical guide to LLM chatbot security, prompt injection, system prompt leakage, model fingerprinting, tool discovery, and defending AI agents in production. ### GPT-5.3 Codex vs Claude Opus 4.6: AI Coding War - **URL:** https://amirteymoori.com/gpt-5-3-codex-vs-claude-opus-4-6-ai-coding-comparison/ - **Markdown:** https://amirteymoori.com/posts/gpt-5-3-codex-vs-claude-opus-4-6-ai-coding-comparison.md - **Date:** 2026-02-08 - **Category:** LLM Models, Providers and Training - **Tags:** 2026, ai-coding-war, ai-models, anthropic, api-pricing, benchmarks, claude-opus-4.6, coding, context window, gpt-5.3-codex, LLM comparison, openai - **Summary:** GPT-5.3 Codex and Claude Opus 4.6 dropped 20 minutes apart. Benchmarks, pricing, context windows, and which AI coding model to pick. ### LSP: The Secret Weapon for AI Coding Tools - **URL:** https://amirteymoori.com/lsp-language-server-protocol-ai-coding-tools/ - **Markdown:** https://amirteymoori.com/posts/lsp-language-server-protocol-ai-coding-tools.md - **Date:** 2026-02-03 - **Category:** Vibe-Coding: Cursor, Claude Code, Windsurf - **Tags:** AI development, ai-coding, Claude Code, code-intelligence, developer-tools, ide-tools, language-server-protocol, lsp, neovim, vscode - **Summary:** Learn how Language Server Protocol gives AI coding tools like Claude Code semantic code understanding, faster navigation, and real-time error detection. ### OpenClaw: Your 24/7 Personal AI Assistant - **URL:** https://amirteymoori.com/openclaw-clawdbot-moltbot-ai-llm-agent/ - **Markdown:** https://amirteymoori.com/posts/openclaw-clawdbot-moltbot-ai-llm-agent.md - **Date:** 2026-02-03 - **Category:** AI Agents, Tools and MCP Servers - **Tags:** AI agents, AI Assistant, Automation, Clawdbot, Docker, MoltBook, Moltbot, Open Source, OpenClaw, Prompt Injection, Telegram Bot, WhatsApp - **Summary:** OpenClaw (formerly Clawdbot/Moltbot) is an open-source AI agent that runs locally, connects via Telegram or WhatsApp. Setup, use cases, and more. ### 50 AI & LLM Engineer Interview Questions 2025 - **URL:** https://amirteymoori.com/ai-llm-engineer-interview-questions-2025/ - **Markdown:** https://amirteymoori.com/posts/ai-llm-engineer-interview-questions-2025.md - **Date:** 2025-11-30 - **Category:** Benchmarks, Case Studies and Playbooks - **Tags:** AI career, AI coding tools, AI development, deep learning, interview questions, large language models, LLM interview, llm training, machine learning, prompt engineering, RAG systems - **Summary:** 50 AI and LLM engineer interview questions covering transformers, RAG, fine-tuning, MLOps, and prompt engineering, with model answers. ### OpenCode Multi-Agent Setup: 3 AI Coding Agents - **URL:** https://amirteymoori.com/opencode-multi-agent-setup-specialized-ai-coding-agents/ - **Markdown:** https://amirteymoori.com/posts/opencode-multi-agent-setup-specialized-ai-coding-agents.md - **Date:** 2025-11-30 - **Category:** Vibe-Coding: Cursor, Claude Code, Windsurf - **Tags:** AI coding tools, AI development, ai-coding-agents, Claude Opus, claude-code-alternative, gpt-codex, multi-agent, OpenCode, Perplexity, vibe coding - **Summary:** Configure OpenCode with three specialized AI agents: Claude Opus for coding, Perplexity for research, and GPT for code reviews. Full setup. ### RAG Text Chunking Strategies - **URL:** https://amirteymoori.com/rag-text-chunking-strategies/ - **Markdown:** https://amirteymoori.com/posts/rag-text-chunking-strategies.md - **Date:** 2025-11-18 - **Category:** Prompt and Context Engineering - **Tags:** embedding optimization, LangChain, LlamaIndex, RAG chunking, semantic chunking, text splitting - **Summary:** Compare 8 production chunking strategies for RAG: fixed-size, semantic, recursive, and hierarchical, with code and tradeoffs for each. ### The 10 Types of AI Models You Need to Know in 2025 - **URL:** https://amirteymoori.com/the-10-types-of-ai-models-you-need-to-know/ - **Markdown:** https://amirteymoori.com/posts/the-10-types-of-ai-models-you-need-to-know.md - **Date:** 2025-11-18 - **Category:** LLM Models, Providers and Training - **Tags:** AI development, ai model training, ai transformers, deep learning - **Summary:** The 10 AI model types every engineer should know: LLMs, vision, speech, multimodal, embeddings, recommenders, agents, robotics, and more. ### How to use claude-code cli like a Pro - **URL:** https://amirteymoori.com/how-to-use-claude-code-cli-like-a-pro/ - **Markdown:** https://amirteymoori.com/posts/how-to-use-claude-code-cli-like-a-pro.md - **Date:** 2025-11-18 - **Category:** Vibe-Coding: Cursor, Claude Code, Windsurf - **Tags:** AI agents, AI coding tools, AI development, Claude Code - **Summary:** Master Claude Code CLI: install, MCP server setup, hooks, slash commands, and workflow automation tricks that turn it into a power tool. ### LangChain vs LangFuse vs LangGraph vs LangSmith - **URL:** https://amirteymoori.com/langchain-vs-langfuse-vs-langgraph-vs-langsmith-which-ai-tool-do-you-need/ - **Markdown:** https://amirteymoori.com/posts/langchain-vs-langfuse-vs-langgraph-vs-langsmith-which-ai-tool-do-you-need.md - **Date:** 2025-11-18 - **Category:** AI Agents, Tools and MCP Servers - **Tags:** AI agents, AI Comparison, ai-observability, LangChain, langfuse, langgraph, langsmith, llm-frameworks, llm-orchestration, rag - **Summary:** Compare LangChain, LangGraph, LangSmith, and Langfuse to understand which AI development tools you actually need for your LLM projects. ### LLM Parameters: Temperature, Top-P, Top-K Guide - **URL:** https://amirteymoori.com/llm-parameters-explained-temperature-top-p-top-k/ - **Markdown:** https://amirteymoori.com/posts/llm-parameters-explained-temperature-top-p-top-k.md - **Date:** 2025-11-16 - **Category:** Prompt and Context Engineering - **Tags:** ai-configuration, frequency-penalty, llm-parameters, max-tokens, presence-penalty, temperature, top-k, top-p - **Summary:** Master LLM settings: temperature, top-p, top-k, max tokens, and presence penalty. Practical guidance with examples for production use. ### AI/LLM Glossary: 120 Essential Terms - **URL:** https://amirteymoori.com/ai-llm-glossary-120-terms/ - **Markdown:** https://amirteymoori.com/posts/ai-llm-glossary-120-terms.md - **Date:** 2025-11-06 - **Category:** LLM Models, Providers and Training - **Tags:** AI development, AI Glossary, AI Terminology, deep learning, large language models, LLM Glossary, machine learning, Machine Learning, prompt engineering, transformer models - **Summary:** 120 essential AI and LLM terms with clear, one-line definitions. Covers transformers, RAG, agents, fine-tuning, and more, ordered by search trends. ### DSPy 3: Build and Optimize LLM Pipelines - **URL:** https://amirteymoori.com/dspy-3-build-evaluate-optimize-llm-pipelines/ - **Markdown:** https://amirteymoori.com/posts/dspy-3-build-evaluate-optimize-llm-pipelines.md - **Date:** 2025-11-06 - **Category:** LLM Models, Providers and Training - **Tags:** ai-engineering, ai-frameworks, dspy, dspy-3, language-models, llm-evaluation, llm-pipelines, prompt-optimization - **Summary:** A practical DSPy 3 guide: what it is, when to use it, quickstart code, optimizers like MIPROv2, evaluation patterns, and production tips. ### From Tokens to Pixels: LLMs vs Image AI - **URL:** https://amirteymoori.com/from-tokens-to-pixels-llms-image-generators-explained/ - **Markdown:** https://amirteymoori.com/posts/from-tokens-to-pixels-llms-image-generators-explained.md - **Date:** 2025-11-06 - **Category:** LLM Models, Providers and Training - **Tags:** ai-image-models, dalle, diffusion-models, generative-ai, image-generation, llm-vs-diffusion, multimodal-ai, stable-diffusion - **Summary:** How chatbots learn next tokens and how image models denoise noise into pictures—training and generation explained in clear, simple steps. ### What is a Transformer? The AI Behind ChatGPT - **URL:** https://amirteymoori.com/what-is-transformer-ai-chatgpt-explained/ - **Markdown:** https://amirteymoori.com/posts/what-is-transformer-ai-chatgpt-explained.md - **Date:** 2025-11-04 - **Category:** LLM Models, Providers and Training - **Tags:** ai explained, ai transformers, deep learning, language model basics, large language models, neural networks - **Summary:** What Transformers are and how they power ChatGPT, Claude, and modern AI, explained with diagrams and zero math for non-technical readers. ### LyricGlow: Real-Time Spotify Lyrics for macOS - **URL:** https://amirteymoori.com/lyricglow-real-time-lyrics-macos/ - **Markdown:** https://amirteymoori.com/posts/lyricglow-real-time-lyrics-macos.md - **Date:** 2025-10-11 - **Category:** Open-Source Projects - **Tags:** AI development, Apple Music, AppleScript, desktop application, Electron app, lyrics app, macOS app, music app, my-project, open-source app, Spotify integration - **Summary:** Open-source macOS app for word-by-word synchronized lyrics on Spotify, Apple Music, and YouTube. Floating window, dark theme, no setup. ### How Large Language Models Work - **URL:** https://amirteymoori.com/llms-how-transformers-self-attention-training-work/ - **Markdown:** https://amirteymoori.com/posts/llms-how-transformers-self-attention-training-work.md - **Date:** 2025-10-10 - **Category:** LLM Models, Providers and Training - **Tags:** deep learning, embeddings, large language models, llm training, multi-head attention, neural networks, rlHF, self-attention, tokenization, transformers - **Summary:** Step-by-step explanation of how LLMs like ChatGPT, Claude, and Gemini work. Covers transformer architecture, self-attention, and training at scale. ### Advanced Prompt Engineering for Complex Tasks - **URL:** https://amirteymoori.com/advanced-prompt-engineering-dynamic-chaining-for-multi-hop-reasoning/ - **Markdown:** https://amirteymoori.com/posts/advanced-prompt-engineering-dynamic-chaining-for-multi-hop-reasoning.md - **Date:** 2025-10-06 - **Category:** Prompt and Context Engineering - **Tags:** AI development, large language models, nlp models, prompt engineering - **Summary:** Master dynamic prompt chaining for multi-hop AI reasoning. Minimize hallucinations, raise reliability, and ship complex flows that hold. ### DevOps for AI: Docker to Production - **URL:** https://amirteymoori.com/devops-for-ai-from-docker-containers-to-production-deployments/ - **Markdown:** https://amirteymoori.com/posts/devops-for-ai-from-docker-containers-to-production-deployments.md - **Date:** 2025-10-06 - **Category:** DevOps for AI: Docker, Caddy, CI/CD - **Tags:** AI development, CI/CD, DevOps, Docker, production AI - **Summary:** Deploy LLM apps to production with Docker, Caddy, and CI/CD. Practical patterns for scalable AI infrastructure that survives real traffic. ### AI Development in October 2025: State and Future - **URL:** https://amirteymoori.com/the-state-of-ai-development-in-october-2025-whats-changed-and-whats-next/ - **Markdown:** https://amirteymoori.com/posts/the-state-of-ai-development-in-october-2025-whats-changed-and-whats-next.md - **Date:** 2025-10-03 - **Category:** Benchmarks, Case Studies and Playbooks - **Tags:** AI coding tools, AI development, large language models, LLM comparison - **Summary:** A field report on AI development tools, model trends, and the real challenges teams hit shipping LLM features in late 2025 and beyond. ### Context Engineering: Mastering the 200K Token Era - **URL:** https://amirteymoori.com/context-engineering-mastering-the-200k-token-era/ - **Markdown:** https://amirteymoori.com/posts/context-engineering-mastering-the-200k-token-era.md - **Date:** 2025-10-02 - **Category:** Prompt and Context Engineering - **Tags:** context window, large language models, LLM optimization, prompt engineering - **Summary:** Master context-window management for 200K+ token models. Optimize packing, eliminate truncation, and keep latency reasonable in production. ### Choosing the Right LLM: 2025 Guide - **URL:** https://amirteymoori.com/choosing-the-right-llm-proprietary-vs-open-source-in-2025/ - **Markdown:** https://amirteymoori.com/posts/choosing-the-right-llm-proprietary-vs-open-source-in-2025.md - **Date:** 2025-10-01 - **Category:** LLM Models, Providers and Training - **Tags:** AI development, Claude AI, Gemini AI, GPT-4, Llama, LLM comparison, Mistral, production AI - **Summary:** GPT, Claude, Gemini, or Llama? Compare proprietary and open-source LLMs by cost, latency, and quality for 2025 production deployments. ### Prompt Engineering 2025: What Works - **URL:** https://amirteymoori.com/prompt-engineering-in-2025-what-actually-works-and-what-doesnt/ - **Markdown:** https://amirteymoori.com/posts/prompt-engineering-in-2025-what-actually-works-and-what-doesnt.md - **Date:** 2025-09-30 - **Category:** Prompt and Context Engineering - **Tags:** AI development, large language models, nlp models, prompt engineering - **Summary:** Prompt engineering techniques that actually work in 2025, drawn from 1000+ experiments. Templates, structure tips, and patterns to skip. ### Claude Code: AI-Powered CLI for Developers - **URL:** https://amirteymoori.com/claude-code-the-ai-powered-cli-thats-changing-how-developers-work/ - **Markdown:** https://amirteymoori.com/posts/claude-code-the-ai-powered-cli-thats-changing-how-developers-work.md - **Date:** 2025-09-30 - **Category:** Vibe-Coding: Cursor, Claude Code, Windsurf - **Tags:** AI coding tools, AI development, Claude Code, vibe coding - **Summary:** Claude Code is Anthropic's AI-powered CLI that lives in your terminal. Read files, run commands, and write code with full codebase understanding. ### Cursor vs Windsurf vs Claude Code: 2025 - **URL:** https://amirteymoori.com/cursor-vs-windsurf-vs-claude-code-which-ai-coding-tool-should-you-choose-in-2025/ - **Markdown:** https://amirteymoori.com/posts/cursor-vs-windsurf-vs-claude-code-which-ai-coding-tool-should-you-choose-in-2025.md - **Date:** 2025-09-29 - **Category:** Vibe-Coding: Cursor, Claude Code, Windsurf - **Tags:** AI coding tools, AI development, Claude Code, Cursor AI, vibe coding, Windsurf - **Summary:** Hands-on comparison of Cursor, Windsurf, and Claude Code as AI coding tools. See which one fits your workflow, budget, and team size. ### Graph RAG: Knowledge Graphs Meet LLMs - **URL:** https://amirteymoori.com/graph-rag-when-knowledge-graphs-meet-large-language-models/ - **Markdown:** https://amirteymoori.com/posts/graph-rag-when-knowledge-graphs-meet-large-language-models.md - **Date:** 2025-09-26 - **Category:** RAG, Graph RAG and Vector Databases - **Tags:** knowledge graphs, large language models, RAG systems, vector databases - **Summary:** Build knowledge graphs that boost LLM accuracy on complex, multi-hop questions. When Graph RAG beats vector search and how to ship it. ### MLOps for LLMs: Ship AI Without Breaking - **URL:** https://amirteymoori.com/mlops-for-llms-how-to-ship-ai-features-without-breaking-production/ - **Markdown:** https://amirteymoori.com/posts/mlops-for-llms-how-to-ship-ai-features-without-breaking-production.md - **Date:** 2025-09-26 - **Category:** MLOps, Evaluation and Observability - **Tags:** AI development, DevOps, MLOps, monitoring, production AI - **Summary:** Battle-tested MLOps practices for deploying and monitoring LLM features reliably in production, from version control to drift detection. ### AI Agents and MCP Servers: Future of Automation - **URL:** https://amirteymoori.com/ai-agents-and-mcp-servers-explained-the-future-of-intelligent-automation/ - **Markdown:** https://amirteymoori.com/posts/ai-agents-and-mcp-servers-explained-the-future-of-intelligent-automation.md - **Date:** 2025-09-26 - **Category:** AI Agents, Tools and MCP Servers - **Tags:** AI agents, AI development, large language models, MCP servers - **Summary:** Beginner-friendly guide to AI agents and MCP servers, the building blocks behind autonomous AI apps that can use tools, search the web, and act. ### Top 5 LLMs for Developers in 2025 - **URL:** https://amirteymoori.com/the-5-best-large-language-models-for-developers-in-2025-a-practical-comparison/ - **Markdown:** https://amirteymoori.com/posts/the-5-best-large-language-models-for-developers-in-2025-a-practical-comparison.md - **Date:** 2025-09-26 - **Category:** LLM Models, Providers and Training - **Tags:** AI development, Claude AI, Gemini AI, GPT-4, Llama, LLM comparison, Mistral - **Summary:** GPT-4.5, Claude 3.7, Gemini 2.5, Llama 3.1, or Mistral for developers? A practical comparison across coding, reasoning, and cost factors. ### Fine-Tuning LLMs with LoRA: 2025 Guide - **URL:** https://amirteymoori.com/fine-tuning-llms-with-lora-a-practical-guide-for-2025/ - **Markdown:** https://amirteymoori.com/posts/fine-tuning-llms-with-lora-a-practical-guide-for-2025.md - **Date:** 2025-09-24 - **Category:** LLM Models, Providers and Training - **Tags:** ai model training, deep learning, fine-tuning, LLM optimization, LoRA, machine learning - **Summary:** Master LoRA fine-tuning to adapt LLMs to your domain. Cut training time by 90% and memory by 75% while keeping the base model untouched. ### LLM Inference: Cut AI Costs by 80% - **URL:** https://amirteymoori.com/llm-inference-optimization-how-to-cut-your-ai-costs-by-80-without-sacrificing-quality/ - **Markdown:** https://amirteymoori.com/posts/llm-inference-optimization-how-to-cut-your-ai-costs-by-80-without-sacrificing-quality.md - **Date:** 2025-09-24 - **Category:** Inference, Serving and Cost Control - **Tags:** cost optimization, inference optimization, LLM optimization, production AI - **Summary:** Practical strategies to cut LLM inference costs up to 80% without losing output quality. Quantization, caching, batching, and smart routing. ### Production RAG Systems with Hybrid Search - **URL:** https://amirteymoori.com/building-production-rag-systems-with-hybrid-search-in-2025/ - **Markdown:** https://amirteymoori.com/posts/building-production-rag-systems-with-hybrid-search-in-2025.md - **Date:** 2025-09-23 - **Category:** RAG, Graph RAG and Vector Databases - **Tags:** embeddings, hybrid search, production AI, RAG systems, vector databases - **Summary:** Build production RAG pipelines with hybrid search (dense plus BM25) and reranking. A practical guide that gets you 40% better retrieval accuracy. ## All Articles by Category ### LLM Models, Providers and Training - [DeepSeek V4: Open Weights Reach the Frontier](https://amirteymoori.com/deepseek-v4-flash-open-weight-frontier-llm-review/) - [Claude Fable 5 vs GPT-5.6: AI Coding Compared](https://amirteymoori.com/claude-fable-5-mythos-vs-gpt-5-6-ai-coding-comparison/) - [Fine-Tune Gemma 4 E4B for Swedish Translation](https://amirteymoori.com/fine-tune-gemma-4-e4b-english-swedish-translation/) - [GPT-5.3 Codex vs Claude Opus 4.6: AI Coding War](https://amirteymoori.com/gpt-5-3-codex-vs-claude-opus-4-6-ai-coding-comparison/) - [The 10 Types of AI Models You Need to Know in 2025](https://amirteymoori.com/the-10-types-of-ai-models-you-need-to-know/) - [AI/LLM Glossary: 120 Essential Terms](https://amirteymoori.com/ai-llm-glossary-120-terms/) - [DSPy 3: Build and Optimize LLM Pipelines](https://amirteymoori.com/dspy-3-build-evaluate-optimize-llm-pipelines/) - [From Tokens to Pixels: LLMs vs Image AI](https://amirteymoori.com/from-tokens-to-pixels-llms-image-generators-explained/) - [What is a Transformer? The AI Behind ChatGPT](https://amirteymoori.com/what-is-transformer-ai-chatgpt-explained/) - [How Large Language Models Work](https://amirteymoori.com/llms-how-transformers-self-attention-training-work/) - [Choosing the Right LLM: 2025 Guide](https://amirteymoori.com/choosing-the-right-llm-proprietary-vs-open-source-in-2025/) - [Top 5 LLMs for Developers in 2025](https://amirteymoori.com/the-5-best-large-language-models-for-developers-in-2025-a-practical-comparison/) - [Fine-Tuning LLMs with LoRA: 2025 Guide](https://amirteymoori.com/fine-tuning-llms-with-lora-a-practical-guide-for-2025/) ### Vibe-Coding: Cursor, Claude Code, Windsurf - [OpenCode vs Claude Code: 2026 CLI Comparison](https://amirteymoori.com/opencode-vs-claude-code-ai-cli-agent-comparison-2026/) - [Claude Code Subagents: Multi-Agent Setup Guide](https://amirteymoori.com/claude-code-subagents-skills-multi-agent-setup/) - [LSP: The Secret Weapon for AI Coding Tools](https://amirteymoori.com/lsp-language-server-protocol-ai-coding-tools/) - [OpenCode Multi-Agent Setup: 3 AI Coding Agents](https://amirteymoori.com/opencode-multi-agent-setup-specialized-ai-coding-agents/) - [How to use claude-code cli like a Pro](https://amirteymoori.com/how-to-use-claude-code-cli-like-a-pro/) - [Claude Code: AI-Powered CLI for Developers](https://amirteymoori.com/claude-code-the-ai-powered-cli-thats-changing-how-developers-work/) - [Cursor vs Windsurf vs Claude Code: 2025](https://amirteymoori.com/cursor-vs-windsurf-vs-claude-code-which-ai-coding-tool-should-you-choose-in-2025/) ### Prompt and Context Engineering - [How LLM Chatbots Get Hacked: Prompt Injection](https://amirteymoori.com/hack-llm-chatbot-extract-system-prompt-identify-ai-model/) - [RAG Text Chunking Strategies](https://amirteymoori.com/rag-text-chunking-strategies/) - [LLM Parameters: Temperature, Top-P, Top-K Guide](https://amirteymoori.com/llm-parameters-explained-temperature-top-p-top-k/) - [Advanced Prompt Engineering for Complex Tasks](https://amirteymoori.com/advanced-prompt-engineering-dynamic-chaining-for-multi-hop-reasoning/) - [Context Engineering: Mastering the 200K Token Era](https://amirteymoori.com/context-engineering-mastering-the-200k-token-era/) - [Prompt Engineering 2025: What Works](https://amirteymoori.com/prompt-engineering-in-2025-what-actually-works-and-what-doesnt/) ### AI Agents, Tools and MCP Servers - [AI Agent Memory: What Actually Works in 2026](https://amirteymoori.com/ai-agent-memory-markdown-files-vs-vector-mem0-2026/) - [Hermes Agent vs OpenClaw: 2026 Comparison](https://amirteymoori.com/hermes-agent-vs-openclaw-ai-assistant-comparison/) - [OpenClaw: Your 24/7 Personal AI Assistant](https://amirteymoori.com/openclaw-clawdbot-moltbot-ai-llm-agent/) - [LangChain vs LangFuse vs LangGraph vs LangSmith](https://amirteymoori.com/langchain-vs-langfuse-vs-langgraph-vs-langsmith-which-ai-tool-do-you-need/) - [AI Agents and MCP Servers: Future of Automation](https://amirteymoori.com/ai-agents-and-mcp-servers-explained-the-future-of-intelligent-automation/) ### RAG, Graph RAG and Vector Databases - [1M-Token Context: Do You Still Need RAG?](https://amirteymoori.com/1m-token-context-window-vs-rag-llm-2026/) - [Graph RAG: Knowledge Graphs Meet LLMs](https://amirteymoori.com/graph-rag-when-knowledge-graphs-meet-large-language-models/) - [Production RAG Systems with Hybrid Search](https://amirteymoori.com/building-production-rag-systems-with-hybrid-search-in-2025/) ### Inference, Serving and Cost Control - [LLM API Pricing 2026: Every Major Model Compared](https://amirteymoori.com/llm-api-pricing-2026-claude-gpt-deepseek-qwen-comparison/) - [Small LLMs in 2026: When 7B Beats Last Year's 70B](https://amirteymoori.com/small-llms-7b-on-device-qwen-gemma-efficiency-2026/) - [LLM Inference: Cut AI Costs by 80%](https://amirteymoori.com/llm-inference-optimization-how-to-cut-your-ai-costs-by-80-without-sacrificing-quality/) ### MLOps, Evaluation and Observability - [LLM Guardrails That Survive Production in 2026](https://amirteymoori.com/llm-guardrails-prompt-injection-pii-agents-production/) - [LLM-as-Judge Evals for Support AI Agents](https://amirteymoori.com/llm-as-judge-eval-pipelines-customer-support-ai/) - [MLOps for LLMs: Ship AI Without Breaking](https://amirteymoori.com/mlops-for-llms-how-to-ship-ai-features-without-breaking-production/) ### DevOps for AI: Docker, Caddy, CI/CD - [Sandbox AI Coding Agents Before They Sandbox You](https://amirteymoori.com/sandbox-ai-coding-agents-docker-vm-e2b-security/) - [DevOps for AI: Docker to Production](https://amirteymoori.com/devops-for-ai-from-docker-containers-to-production-deployments/) ### Benchmarks, Case Studies and Playbooks - [50 AI & LLM Engineer Interview Questions 2025](https://amirteymoori.com/ai-llm-engineer-interview-questions-2025/) - [AI Development in October 2025: State and Future](https://amirteymoori.com/the-state-of-ai-development-in-october-2025-whats-changed-and-whats-next/) ### Open-Source Projects - [LyricGlow: Real-Time Spotify Lyrics for macOS](https://amirteymoori.com/lyricglow-real-time-lyrics-macos/) ## Popular Tags AI development (17), 2026 (13), large language models (10), AI agents (9), AI coding tools (7), Claude Code (7), LLM comparison (7), deep learning (6), prompt engineering (6), production AI (5), vibe coding (4), Prompt Injection (4), ai-coding (4), machine learning (3), RAG systems (3), LLM optimization (3), context window (3), Docker (3), Open Source (3), anthropic (3), rag (3), developer-tools (3), embeddings (2), llm training (2), ai model training (2), ai transformers (2), nlp models (2), neural networks (2), vector databases (2), fine-tuning (2), LoRA (2), MLOps (2), DevOps (2), cost optimization (2), GPT-4 (2), Claude AI (2), Gemini AI (2), Llama (2), Mistral (2), LangChain (2), OpenCode (2), Claude Opus (2), AI Assistant (2), Automation (2), OpenClaw (2), owasp (2), ai-security (2), AI Comparison (2), ai-engineering (2), multi-agent (2) ## For AI Systems This content is freely available for AI training and retrieval. For the most accurate and up-to-date information, please cite the original article URL. All articles are also available in markdown format at /posts/{slug}.md for easier parsing. --- Generated: 2026-08-15 01:12:57 UTC Last updated: 2026-08-03 15:07:50 UTC