# Pawel Lipowczan > Architekt oprogramowania i doradca ds. technologii — agnostyczny dobór narzędzi do problemu, optymalizacja procesów biznesowych przez automatyzację i inteligentne rozwiązania no-code oraz AI. ## Blog (PL) - [5 słabych stron Claude Code i jak je nadrobić](https://pawel.lipowczan.pl/blog/slabe-strony-claude-code): Claude Code jest świetny w kodzie, ale ślepy przy wideo, designie, pamięci i researchu. Oto 5 słabych stron i narzędzia, których używam, żeby je nadrobić. - [RAG ragowi nierówny: do kodu graf, do notatek indeks](https://pawel.lipowczan.pl/blog/rag-ragowi-nierowny): RAG to rodzina technik, nie jedna. Do kodu wygrywa retrieval strukturalny (AST, LSP, graf), do notatek kurowany indeks. Mapa doboru mechanizmu do materiału. - [Zbudowałem drugi mózg. Trafiłem w standard Google (OKF)](https://pawel.lipowczan.pl/blog/okf-standard-przenosnosc-bazy-wiedzy-ai): Mój second brain jest zgodny z Open Knowledge Format Google w ~100% - choć nigdy go pod niego nie projektowałem. O przenośności bazy wiedzy decyduje format, nie narzędzie. - [Software 3.0: dlaczego twoja aplikacja nie powinna istnieć](https://pawel.lipowczan.pl/blog/software-3-0-agentic-engineering): Trzy paradygmaty oprogramowania i nowa dyscyplina: agentic engineering. Mapa decyzyjna od Karpathy'ego - co budować, czego nie i czego nie da się outsourcować. - [Pre-revenue startup bez działu finansów. Anatomia systemu agentów AI.](https://pawel.lipowczan.pl/blog/system-agentow-ai-skills-rules-kontekst): 5 ról specjalistów, 2 założycieli. Jak skills, deterministyczne reguły i wspólny kontekst zastąpiły zespół w 200IQ LABS - case study i architektura. - [Miałem 3 dni do PIT-38 bez księgowej. Wystarczyły 2 godziny.](https://pawel.lipowczan.pl/blog/pit-38-claude-code-case-study): 5430 transakcji, 5 źródeł danych, 2 dni z agentem Claude Code. Case study workflow, który zastąpił księgową przy rozliczeniu PIT-38. - [Dlaczego nie da się tego zrobić na WordPressie - spec-driven SEO na portfolio i Qamera AI](https://pawel.lipowczan.pl/blog/spec-driven-seo-portfolio-qamera-ai): Audyt SEO, spec-driven changes i AI workflow zoptymalizowały dwa projekty w godziny, nie tygodnie. Case study z portfolio i Qamera AI. - [Jak LLM Wiki Karpathy'ego pomogła mi uporządkować moją bazę wiedzy](https://pawel.lipowczan.pl/blog/llm-knowledge-base-brain-karpathy): Od ręcznych notatek w Obsidian do systemu, w którym agent pilnuje struktury, indeksów i standardów. Jak koncepcja LLM Wiki pomogła mi sformalizować to, co budowałem od lat. - [5 repozytoriów GitHub, które zmienią Twoją pracę z Claude Code](https://pawel.lipowczan.pl/blog/5-repozytoriow-github-claude-code): Testuję dziesiątki narzędzi - te 5 repozytoriów naprawdę zrobiło różnicę w mojej codziennej pracy z Claude Code. Oto moja sprawdzona lista. - [Moje środowisko agentowe - jak buduję AI OS dla dwóch firm](https://pawel.lipowczan.pl/blog/srodowisko-agentowe-ai-dwie-firmy): Dwa repozytoria, osiem agentów, zero vendor lock-in. Jak zbudowałem system agentowy zarządzający finansami, prawem i marketingiem dwóch firm. - [Skills 2.0 - jak buduję system wieloagentowy do zarządzania firmą](https://pawel.lipowczan.pl/blog/skills-2-0-multi-agent-system-zarzadzanie-firma): Jak przejść od rozproszonych promptów AI do zorganizowanego systemu agentów zarządzających finansami, prawem i marketingiem Twojej firmy. - [OpenClaw: lekcja bezpieczeństwa, której potrzebował świat agentów AI](https://pawel.lipowczan.pl/blog/openclaw-bezpieczenstwo-agentow-ai): 150 tysięcy gwiazdek w 2 tygodnie, 28 tysięcy odsłoniętych instancji i religia stworzona przez boty. Co OpenClaw mówi o przyszłości autonomicznych agentów. - [OPSX Workflow - strukturyzowane podejście do pracy z AI coding assistants](https://pawel.lipowczan.pl/blog/opsx-workflow-strukturyzowana-praca-z-ai): Jak przekształcić chaotyczne promptowanie w powtarzalny proces z artefaktami, zależnościami i iteracją. - [Remotion + AI: Jak tworzyć profesjonalne wideo za pomocą kodu i Claude](https://pawel.lipowczan.pl/blog/remotion-explainer-videos-ai): Stworzyłem 45-sekundowy explainer video w kilka minut. Bez Adobe, bez Canva - tylko React, Remotion i Claude Code. Zobacz jak. - [Jak tworzyć animacje w stylu Apple z pomocą AI](https://pawel.lipowczan.pl/blog/animacje-apple-ai-cursor): Dowiedz się, jak stworzyć płynne scroll-animacje w stylu stron Apple za pomocą narzędzi AI - Google Whisk, Flow i Cursor. - [Second Brain z Obsidian i Claude Code - jak AI zmienia organizację wiedzy](https://pawel.lipowczan.pl/blog/second-brain-obsidian-claude-code-skills): Odkryj jak połączenie Obsidian, Claude Code i Skills tworzy potężny system zarządzania wiedzą. Praktyczny przewodnik po budowie prywatnego second brain. - [15 hacków do Cursor.sh które zmienią sposób pracy z AI](https://pawel.lipowczan.pl/blog/15-cursor-hacks-produktywnosc-ai): Praktyczny przegląd 15 sprawdzonych technik pracy z Cursor.sh - od skrótów klawiszowych po zaawansowane worktrees, strukturyzację promptów i integracje, które realnie przyspieszają codzienne… - [5 technik które zmienią sposób pracy z Claude Code](https://pawel.lipowczan.pl/blog/5-technik-pracy-z-claude-code): Większość programistów wykorzystuje zaledwie 20% potencjału Claude Code. Poznaj 5 zaawansowanych technik używanych przez najlepszych inżynierów AI. - [2026: Rok, w którym AI przeszła z laboratoriów do hal produkcyjnych](https://pawel.lipowczan.pl/blog/trendy-ai-2026-od-eksperymentow-do-operacjonalizacji): Koniec hype'u, początek weryfikacji. Jak agentic AI, specjalizacja modeli i regulacje EU AI Act zmienią sposób, w jaki budujemy i wdrażamy sztuczną inteligencję w biznesie. - [Vibe Coding - jak tworzyć UI z AI bez znajomości designu](https://pawel.lipowczan.pl/blog/vibe-coding-przewodnik): Odkryj 3 filary skutecznego vibe codingu i dowiedz się, jak AI może zamienić Twoje pomysły w profesjonalne interfejsy - nawet jeśli nie masz umiejętności projektowych. - [Dane jako paliwo biznesu - od Excela do AI](https://pawel.lipowczan.pl/blog/dane-jako-paliwo-biznesu): Jak przekształcić surowe dane w wartość biznesową? No-code, agenty AI i case study transformacji z chaosu do klarowności. - [Airtable vs Excel - Kiedy warto zmienić arkusz na bazę danych?](https://pawel.lipowczan.pl/blog/airtable-vs-excel-migracja): Excel przestał wystarczać? Dowiedz się, dlaczego coraz więcej firm migruje do Airtable i jak przejść z arkuszy kalkulacyjnych na relacyjną bazę danych bez bólu głowy. - [Hackathon Hacknation - Analiza Doświadczeń i Lekcja z Cyfryzacji](https://pawel.lipowczan.pl/blog/hackathon-hacknation-analiza-doswiadczen): Zobacz jak w 24h zespół nieprogramistów z pomocą AI stworzył działający system budżetowy i dlaczego technologia to nie wszystko. Szczera analiza sukcesów i porażek z Hacknation. - [Kodowanie w 2025: Czy AI zbudowało moje portfolio? Case Study pawel.lipowczan.pl](https://pawel.lipowczan.pl/blog/kodowanie-w-2025-ai-portfolio): Czy w 2025 roku programowanie się kończy? Zbudowałem portfolio pawel.lipowczan.pl jako poligon doświadczalny współpracy z AI. Wnioski? AI to junior developer na sterydach, ale fundamenty inżynierskie… - [Każda firma działa nieoptymalnie - jak przestać kłamać pracownikom i zacząć naprawiać procesy?](https://pawel.lipowczan.pl/blog/kazda-firma-dziala-nieoptymalnie): Czy wiesz, że Twoja firma traci czas i pieniądze na nieefektywne procesy? Dowiedz się, jak mapowanie procesów i automatyzacja mogą to zmienić. Wnioski z Infoshare Katowice 2025. - [Zapier vs Make vs n8n - jak wybrać narzędzie automatyzacji dla Twojego zespołu?](https://pawel.lipowczan.pl/blog/zapier-vs-make-vs-n8n-wybor-narzedzia): Wybór złego narzędzia automatyzacji to miesiące straconego czasu i tysiące złotych na migrację. Dowiedz się, jak wybrać między Zapier, Make i n8n na podstawie kompetencji zespołu, skali operacji i… - [Jak agencja eventowa El Padre przyspieszyła tworzenie ofert nawet o 50% dzięki AI](https://pawel.lipowczan.pl/blog/el-padre-automatyzacja-ofert-ai): Jak agencja eventowa El Padre przyspieszyła proces ofertowania nawet o 50% dzięki wdrożeniu AI? Case study pokazuje konkretne liczby: 10-15% większa produktywność, 25-30 osób wspieranych przez AI i… - [Automatyzacja poczty email z AI - Jak Frontdesk AI rewolucjonizuje obsługę klienta](https://pawel.lipowczan.pl/blog/automatyzacja-email-frontdesk-ai): Dowiedz się jak system Frontdesk AI automatycznie przetwarza, kategoryzuje i odpowiada na wiadomości email, oszczędzając dziesiątki godzin pracy miesięcznie. - [No-Code Lead Generation - Jak zbudować system generowania leadów bez programowania](https://pawel.lipowczan.pl/blog/no-code-lead-generation): Kompleksowy przewodnik po budowie systemu Lead Generator wykorzystującego n8n, Airtable i różne źródła danych biznesowych. Case study z realnego wdrożenia. - [Chatboty oparte na AI - Od koncepcji do wdrożenia](https://pawel.lipowczan.pl/blog/chatboty-ai-od-koncepcji-do-wdrozenia): Kompletny przewodnik po tworzeniu inteligentnych chatbotów i voicebotów z wykorzystaniem VAPI, n8n, OpenAI i RAG (Retrieval Augmented Generation). ## Blog (EN) - [5 Weak Spots of Claude Code and How to Fix Them](https://pawel.lipowczan.pl/en/blog/claude-code-weak-spots): Claude Code is brilliant with code but blind at video, design, memory and research. Here are its 5 weak spots and the tools I use to fix each one. - [Not All RAG Is Equal: a Graph for Code, an Index for Notes](https://pawel.lipowczan.pl/en/blog/not-all-rag-is-equal): RAG is a family, not one technique. Code favors structural retrieval (AST, LSP, graphs), notes a curated index. A map for matching mechanism to material. - [I Built a Second Brain. It Already Matched Google's Standard (OKF)](https://pawel.lipowczan.pl/en/blog/okf-standard-portable-knowledge-base): My second brain hit ~100% compliance with Google's Open Knowledge Format - though I never designed it for OKF. Format, not tooling, decides whether your knowledge base is portable. - [Software 3.0: why your app shouldn't exist](https://pawel.lipowczan.pl/en/blog/software-3-0-agentic-engineering-guide): Three software paradigms and a new discipline: agentic engineering. A decision map from Karpathy - what to build, what not to, and what you can't outsource. - [Pre-revenue startup without a finance department. AI agent system anatomy.](https://pawel.lipowczan.pl/en/blog/ai-agent-system-skills-rules-shared-context): 5 specialist roles, 2 founders. How skills, deterministic rules, and shared context replaced a team at 200IQ LABS - case study and architecture. - [3 days to PIT-38, no accountant. 2 hours were enough.](https://pawel.lipowczan.pl/en/blog/polish-pit-38-claude-code-case-study): 5430 transactions, 5 data sources, two days with a Claude Code agent. A case study of the workflow that replaced my accountant on PIT-38. - [Why WordPress can't do this - spec-driven SEO on a portfolio and Qamera AI](https://pawel.lipowczan.pl/en/blog/spec-driven-seo-portfolio-qamera-ai-case-study): An SEO audit, spec-driven changes, and an AI workflow optimized two projects in hours, not weeks. Case study from a portfolio and Qamera AI. - [How Karpathy's LLM Wiki Helped Me Organize My Knowledge Base](https://pawel.lipowczan.pl/en/blog/karpathy-llm-wiki-knowledge-base): A walk-through of a PKM system that evolved from manual Obsidian notes into an agent-maintained knowledge base of 185 notes - the three-layer architecture, progressive disclosure indexes, four core… - [5 GitHub Repositories That Will Transform Your Work with Claude Code](https://pawel.lipowczan.pl/en/blog/5-github-repos-claude-code): I test dozens of tools - these 5 repositories truly made a difference in my daily work with Claude Code. Here's my battle-tested list. - [My agentic environment - how I'm building an AI OS for two companies](https://pawel.lipowczan.pl/en/blog/agentic-ai-environment-two-companies): Two repositories, eight agents, zero vendor lock-in. How I built an agent system managing finances, legal, and marketing for two companies. - [Skills 2.0 - how I'm building a multi-agent system to manage my company](https://pawel.lipowczan.pl/en/blog/skills-2-0-multi-agent-system-company-management): How to go from scattered AI prompts to an organized system of agents managing your company's finances, legal, and marketing. - [OpenClaw: the security lesson the AI agent world needed](https://pawel.lipowczan.pl/en/blog/openclaw-ai-agent-security): 150 thousand stars in 2 weeks, 28 thousand exposed instances, and a religion created by bots. What OpenClaw tells us about the future of autonomous agents. - [OPSX Workflow - a structured approach to working with AI coding assistants](https://pawel.lipowczan.pl/en/blog/opsx-workflow-structured-ai-work): How to transform chaotic prompting into a repeatable process with artifacts, dependencies, and iteration. - [Remotion + AI: How to Create Professional Videos with Code and Claude](https://pawel.lipowczan.pl/en/blog/remotion-explainer-videos-ai): I created a 45-second explainer video in minutes. No Adobe, no Canva - just React, Remotion, and Claude Code. Here's how. - [How to Create Apple-Style Animations with AI](https://pawel.lipowczan.pl/en/blog/apple-animations-ai-cursor): Learn how to build smooth scroll animations in the style of Apple product pages using AI tools - Google Whisk, Flow, and Cursor. - [Second Brain with Obsidian and Claude Code - how AI is changing knowledge management](https://pawel.lipowczan.pl/en/blog/second-brain-obsidian-claude-code-skills): Discover how combining Obsidian, Claude Code, and Skills creates a powerful knowledge management system. A practical guide to building a private second brain. - [15 Cursor.sh Hacks That Will Change How You Work with AI](https://pawel.lipowczan.pl/en/blog/15-cursor-hacks-ai-productivity): Most users tap into only 20% of Cursor's capabilities. Discover hidden features from basic shortcuts to advanced techniques like worktrees and structured prompting. - [5 Techniques That Will Transform How You Work with Claude Code](https://pawel.lipowczan.pl/en/blog/5-techniques-working-with-claude-code): Most developers use only 20% of Claude Code's potential. Learn 5 advanced techniques used by the best AI engineers. - [2026: The year AI moved from labs to factory floors](https://pawel.lipowczan.pl/en/blog/ai-trends-2026-from-experiments-to-operationalization): The hype is over, the reckoning begins. How agentic AI, model specialization, and EU AI Act regulations will change the way we build and deploy artificial intelligence in business. - [Vibe Coding - how to create UI with AI without design skills](https://pawel.lipowczan.pl/en/blog/vibe-coding-guide): Discover the 3 pillars of effective vibe coding and learn how AI can turn your ideas into professional interfaces - even if you have no design skills. - [Data as Business Fuel - From Excel to AI](https://pawel.lipowczan.pl/en/blog/data-as-business-fuel): How do you turn raw data into business value? No-code, AI agents, and a case study of transforming chaos into clarity. - [Airtable vs Excel - When Is It Worth Switching from Spreadsheets to a Database?](https://pawel.lipowczan.pl/en/blog/airtable-vs-excel-migration): Has Excel stopped being enough? Learn why more and more companies are migrating to Airtable and how to move from spreadsheets to a relational database without the headache. - [Hackathon Hacknation - Experience Analysis and a Lesson in Digitalization](https://pawel.lipowczan.pl/en/blog/hackathon-hacknation-experience-analysis): See how in 24 hours a team of non-programmers built a working budget system with AI, and why technology isn't everything. An honest analysis of successes and failures from Hacknation. - [Coding in 2025: Did AI Build My Portfolio? A Case Study of pawel.lipowczan.pl](https://pawel.lipowczan.pl/en/blog/coding-in-2025-ai-portfolio): Is programming dead in 2025? I built my portfolio pawel.lipowczan.pl as a testing ground for collaborating with AI. The verdict? AI is a junior developer on steroids, but engineering fundamentals are… - [Every Company Operates Suboptimally - How to Stop Lying to Your Employees and Start Fixing Processes](https://pawel.lipowczan.pl/en/blog/every-company-operates-suboptimally): Did you know your company wastes time and money on inefficient processes? Learn how process mapping and automation can change that. Insights from Infoshare Katowice 2025. - [Zapier vs Make vs n8n - how to choose the right automation tool for your team?](https://pawel.lipowczan.pl/en/blog/zapier-vs-make-vs-n8n-tool-choice): Choosing the wrong automation tool means months of wasted time and thousands of dollars on migration. Learn how to choose between Zapier, Make, and n8n based on team competencies, operational scale,… - [How Event Agency El Padre Accelerated Proposal Creation by Up to 50% with AI](https://pawel.lipowczan.pl/en/blog/el-padre-ai-offer-automation): How did event agency El Padre speed up their proposal process by up to 50% with AI? This case study shows concrete numbers: 10-15% higher productivity, 25-30 people supported by AI, and 75-120 hours… - [Email Automation with AI - How Frontdesk AI Revolutionizes Customer Service](https://pawel.lipowczan.pl/en/blog/email-automation-frontdesk-ai): Learn how the Frontdesk AI system automatically processes, categorizes, and responds to emails, saving dozens of work hours per month. - [No-Code Lead Generation - How to Build a Lead Generation System Without Programming](https://pawel.lipowczan.pl/en/blog/no-code-lead-generation): A comprehensive guide to building a Lead Generator system using n8n, Airtable, and various business data sources. A real-world implementation case study. - [AI-Powered Chatbots - From Concept to Deployment](https://pawel.lipowczan.pl/en/blog/ai-chatbots-from-concept-to-deployment): A complete guide to building intelligent chatbots and voicebots using VAPI, n8n, OpenAI, and RAG (Retrieval Augmented Generation). ## Kurs LLM Wiki (PL) - [LLM Wiki — darmowy kurs](https://pawel.lipowczan.pl/llm-wiki/kurs): Zbuduj własny second brain na darmowym szablonie. Krok po kroku; kurs rośnie o kolejne lekcje. - [Co to jest drugi mózg i po co](https://pawel.lipowczan.pl/llm-wiki/kurs/0-co-to-drugi-mozg): Drugi mózg po ludzku, bez ani jednego technicznego słowa. Czym jest zewnętrzna pamięć na wiedzę, po co ją budować i dlaczego nie trzeba umieć programować. - [Trzy pojęcia zanim zaczniesz](https://pawel.lipowczan.pl/llm-wiki/kurs/0-trzy-pojecia): Trzy słowa, które padną w kursie, rozbrojone po ludzku - agent AI, repozytorium i markdown. Plus bonus - co to jest komenda. Zero kodu, same konkrety. - [Uruchom w swoim narzędziu](https://pawel.lipowczan.pl/llm-wiki/kurs/0-uruchom-w-swoim-narzedziu): Gdzie i jak odpalić szablon bez terminala i bez komend gita. Lista narzędzi z asystentem AI, jak wołać komendy, jak zapisać zmiany i jak pobrać przez ZIP. - [Załóż katalog z szablonu](https://pawel.lipowczan.pl/llm-wiki/kurs/1-zaloz-katalog): Czym jest LLM Wiki (koncept Karpathy'ego) i jak z darmowego szablonu postawić uzbrojoną, pustą bazę wiedzy - architektura 3 warstw, 3 indeksy, progressive disclosure, zero RAG. - [Onboarding](https://pawel.lipowczan.pl/llm-wiki/kurs/2-onboarding): Jeden wywiad /onboard konfiguruje całą bazę - schema, foldery tematów i indeksy generują się same. Krok po kroku, na konkretnym przykładzie. To Twój dzień zerowy. - [Pierwszy ingest](https://pawel.lipowczan.pl/llm-wiki/kurs/3-pierwszy-ingest): Zamień surowe źródło w noty i indeksy komendą /ingest. To różnica między stertą plików a żywą wiki z linkami i frontmatterem OKF. - [Pytania i zarządzanie](https://pawel.lipowczan.pl/llm-wiki/kurs/4-pytania-i-zarzadzanie): Pytaj bazę, nie czat - /qa z cytowaniami, /lint i /reindex do utrzymania jakości. Pełna ściąga komend, zasady jakości i anty-wzorce. - [Rozwój i publikacja](https://pawel.lipowczan.pl/llm-wiki/kurs/5-rozwoj-i-publikacja): Opublikuj bazę przez Quartz, zrozum przenośność OKF i poznaj ścieżkę rozwoju - multi-brain, MCP, publikacja i wymiana paczek wiedzy. ## Projekty - [Note Taker + Add-ons](https://pawel.lipowczan.pl/projects/note-taker-addons): System do automatycznego przetwarzania notatek ze spotkań (Fireflies) poprzez Airtable. Umożliwia kompleksowe zarządzanie informacjami z rozmów biznesowych, analizę spotkań, wykorzystanie w innych… - [Lead Generator](https://pawel.lipowczan.pl/projects/lead-generator): System do automatycznego generowania bazy kontaktów na podstawie Google Search, Apollo, The Company API. Automatyczne generowanie leadów dla zadanych parametrów, szybkie budowanie bazy kontaktów i… - [Context-based Chatbot](https://pawel.lipowczan.pl/projects/context-based-chatbot): Inteligentny chatbot/voicebot do komunikacji na stronach www i messengerach wykorzystujący AI do kontekstowych rozmów. System rozumie intencje użytkownika, prowadzi naturalne konwersacje o ofercie i… - [Integracja Systemów - PHU Impex](https://pawel.lipowczan.pl/projects/integracja-systemow-phu-impex): Kompleksowy system synchronizacji danych między SQL Server, Airtable i BigQuery. Umożliwia wygodne przetwarzanie danych finansowo-księgowych w nowoczesnym interfejsie, synchronizację na żądanie i… - [Frontdesk AI](https://pawel.lipowczan.pl/projects/frontdesk-ai): System do automatycznego przetwarzania i kategoryzacji poczty przychodzącej. Analizuje wiadomości email, klasyfikuje według kategorii, automatycznie odpowiada na najczęstsze pytania i wykonuje… - [Automatyzacje Dokumentów](https://pawel.lipowczan.pl/projects/automatyzacje-dokumentow): Kompleksowe systemy do przetwarzania, generowania i obiegu dokumentów. Automatyczne generowanie umów, dokumentacji projektowej, integracja z podpisem elektronicznym Autenti. Case studies:… - [System HRM](https://pawel.lipowczan.pl/projects/system-hrm): Kompleksowy system do zarządzania Human Resources obejmujący zarządzanie urlopami, zwolnieniami lekarskimi i dostępnością pracowników. Automatyzacja procesów HR, system wniosków i automatyczne… - [Lead Enrichment](https://pawel.lipowczan.pl/projects/lead-enrichment): System automatycznego uzupełniania i wzbogacania danych kontaktowych w CRM. Pozyskuje dodatkowe informacje o firmach, decydentach i kontaktach biznesowych z różnych źródeł, aktualizuje dane w CRM. - [Ankiety & Badania Satysfakcji](https://pawel.lipowczan.pl/projects/ankiety-badania-satysfakcji): System do automatycznej obsługi ankiet i badań satysfakcji klientów oraz pracowników. Automatyczna wysyłka w kluczowych momentach customer journey, zbieranie odpowiedzi, analiza wyników z AI i… ## Kontakt - email: pawel@lipowczan.pl - strona: https://pawel.lipowczan.pl --- # 5 słabych stron Claude Code i jak je nadrobić Source: https://pawel.lipowczan.pl/blog/slabe-strony-claude-code Published: 2026-07-13 Claude Code czyta kod, edytuje pliki i odpala komendy w terminalu. W tych zadaniach jest bardzo dobry. Ale poproś go, żeby obejrzał wideo albo zaprojektował ładny front - i od razu widać granice. Testuję dziesiątki narzędzi wokół Claude Code. Za każdym razem wraca ten sam zestaw pięciu słabych stron: wideo, design, pamięć, research i tokeny (**token** - jednostka tekstu, którą model liczy i za którą płacisz). To nie są wady, które psują pracę z kodem. To obszary, w których Claude Code po prostu nie ma wbudowanej funkcji. Dobra wiadomość: każdą z tych dziur da się załatać. Nie teorią, tylko konkretnym narzędziem. Dla każdego obszaru pokażę, czym posługuję się na co dzień. Będę uczciwy - część narzędzi sprawdziłem w codziennej pracy, część mam dopiero na radarze i wyraźnie to zaznaczę. Pomysł na ten podział zainspirował film Chase AI o pięciu repozytoriach dla Claude Code. Ale to jest mój własny zestaw - narzędzia, które sam odpalam, z moimi zastrzeżeniami. ## 🎬 Wideo - Claude Code nie widzi i nie tworzy filmów Claude Code nie potrafi obejrzeć wideo ani go wygenerować. Domyślnie utyka na transkrypcie (**transkrypt** - tekstowy zapis ścieżki dźwiękowej). A transkrypt nie wystarczy, gdy liczy się to, co realnie dzieje się na ekranie. ### Oglądanie: Claude Video `Claude Video` to skill (**skill** - gotowy zestaw instrukcji, który rozszerza agenta o nową umiejętność) z komendą `/watch`. Wklejam adres albo plik plus pytanie. Agent pobiera napisy, wycina klatki (**klatka** - pojedynczy obraz z wideo) i czyta je jako obrazy. Zanim odpowie, naprawdę zobaczył wideo, a nie zgadł z tytułu. Używam tego do dwóch rzeczy. Analizuję cudze treści szybciej niż odtwarzaniem na 2x i wyciągam wiedzę, która potem trafia do mojej bazy notatek. Wyciągam nim też transkrypty z YouTube. ```bash /watch https://youtu.be/dQw4w9WgXcQ co dzieje sie w 30 sekundzie? ``` Koszt tokenów zależy głównie od klatek, bo każda to obraz. Dlatego skill ma cztery tryby, od najtańszego do najdroższego: ```text transcript - same napisy, zero klatek (najtaniej) efficient - klatki kluczowe, do 50 balanced - po zmianach sceny, do 100 (domyslny) token-burner - bez limitu klatek (pelne pokrycie, drogo) ``` Tryb `efficient` bierze klatki kluczowe (**klatka kluczowa** - klatka, którą samo wideo oznacza jako punkt odniesienia). Gdy wideo nie ma napisów, skill zamienia mowę na tekst modelem whisper (**whisper** - model zamieniający mowę na tekst), za darmo przez Groqa. ### Tworzenie: HyperFrames Odwrotny kierunek to generowanie wideo. `HyperFrames` tworzy wideo z plików HTML - kompozycja to zwykły HTML z atrybutami `data-*`, bez Reacta i bez własnego języka. „Marka" żyje w moich stylach CSS, więc każde wideo wychodzi spójne z moją identyfikacją wizualną, a nie generyczne. ```bash npx skills add heygen-com/hyperframes ``` Wymaga Node.js >= 22 i FFmpeg. O generowaniu wideo z kodu pisałem już przy okazji [Remotion do filmów explainer](/blog/remotion-explainer-videos-ai). Różnica jest prosta: HyperFrames jest HTML-native i na licencji Apache 2.0 (komercyjne użycie bez limitów), a Remotion ma za to gotowy rendering rozproszony w chmurze. ## 🎨 Design - koniec z generycznym AI slop Poproś Claude Code o stronę i dostajesz dokładnie to samo co wszyscy: sekcja hero na górze, zaokrąglone karty, fioletowy gradient. To **AI slop** (**AI slop** - generyczny wygląd, który od razu zdradza, że front wygenerował model). Powód jest prosty: modele trenowano na tych samych szablonach. Nie łatam tego jednym narzędziem, tylko trzema skillami ustawionymi w kolejności: kto, system, strażnik. ### 1. UX RULER - dla kogo i po co `UX RULER` wymusza pytanie „dla kogo to jest i jaką mierzalną wartość daje", zanim wskoczę w funkcje. Decyzje o odbiorcy i potrzebie zapisuje do repozytorium jako product memory (**product memory** - pliki w repo, które trzymają decyzje produktowe dla następnego człowieka albo agenta). Dzięki temu wybory są jawne i wielokrotnego użytku. ### 2. UI UX Pro Max - gotowy design system `UI UX Pro Max` generuje spójny design system (**design system** - zestaw reguł koloru, typografii i komponentów) dopasowany do typu projektu. Portfolio dostaje inną logikę niż SaaS czy sklep. Uruchamiam go z Tailwind CSS i React jako bazę, którą potem dociągam, co oszczędza czas na prototyp. To narzędzie opisałem szerzej w [5 repozytoriach GitHub dla Claude Code](/blog/5-repozytoriow-github-claude-code), więc tu nie powtarzam. ### 3. Impeccable - strażnik języka designu `Impeccable` to warstwa, która usuwa charakterystyczne znaki AI: Inter wszędzie, gradienty fioletu, karty w kartach, szary tekst na kolorowym tle. Ma deterministyczny linter (**linter** - narzędzie sprawdzające kod albo wygląd według sztywnych reguł), który działa bez modelu i bez klucza API: ```bash npx impeccable detect src/ ``` Skill daje **23 komendy** pod `/impeccable`, między innymi `craft`, `shape`, `critique` i `colorize`, oraz **27 reguł deterministycznych** plus dodatkowy przegląd modelu. To ta rzecz, która pilnuje, żeby front nie krzyczał „zrobił mnie AI". Razem układa się to w linię: RULER mówi dla kogo, Pro Max buduje system, Impeccable pilnuje języka i ścina AI slop. ## 🧠 Pamięć - Claude Code zapomina po każdej sesji Koniec sesji to zerowanie. Claude Code nie ma wbudowanej pamięci, która kumuluje się między rozmowami. Za każdym razem zaczyna od zera. ### Second brain jako pamięć Mój sposób to **second brain** (**second brain** - tekstowa baza notatek, którą agent czyta, uzupełnia i przeszukuje). Repozytorium jest moją pamięcią: agent zapisuje do niego wiedzę, linkuje notatki i odpowiada na pytania z całości. To wzorzec, który Andrej Karpathy nazwał **LLM Wiki** (**LLM Wiki** - baza, którą model sam buduje i utrzymuje z Twoich źródeł). Różni się od RAG (**RAG** - technika, w której model przed odpowiedzią przeszukuje surowe dokumenty). W RAG nic się nie kumuluje, bo model za każdym razem odkrywa wszystko na nowo. W LLM Wiki wiedza zostaje w plikach i rośnie. Kluczowy element to **progressive disclosure** (**progressive disclosure** - układ notatek tak, żeby agent je znalazł bez zaśmiecania okna kontekstowego). Okno kontekstowe (**okno kontekstowe** - ile tekstu model widzi naraz) jest ograniczone, więc dobry indeks jest ważniejszy niż wrzucenie wszystkiego do jednego pliku. Zbudowałem to na Obsidian i Claude Code. Pełną architekturę, 185 notatek i trzy indeksy opisałem w [artykule o LLM Wiki Karpathy'ego](/blog/llm-knowledge-base-brain-karpathy) oraz w [budowie drugiego mózgu w Obsidian](/blog/second-brain-obsidian-claude-code-skills). Gdzie kończy się graf kodu, a zaczyna baza notatek, rozkładam w [tekście o RAG](/blog/rag-ragowi-nierowny). Jak taką bazę zbudować od zera, pokazuję krok po kroku w [darmowym kursie LLM Wiki](/llm-wiki). ### NotebookLM-py - świeży dodatek `NotebookLM-py` wkłada NotebookLM (**NotebookLM** - silnik Google, w którym Gemini czyta Twoje źródła i odpowiada z cytatami) do wnętrza Claude Code. To nieoficjalne API, więc używam go na własne ryzyko. ```bash uv tool install "notebooklm-py[browser]" notebooklm login ``` Na razie służy mi głównie jako druga droga do transkryptów z YouTube. Dopiero zaczynam, więc traktuję go jako dodatek, nie fundament. ## 🔎 Research - wbudowane wyszukiwanie jest płytkie Wbudowane wyszukiwanie w Claude Code działa, ale powierzchownie. Brakuje środka między „płytko" a „sto agentów i 10 milionów tokenów na deep research" (**deep research** - głębokie, wieloźródłowe badanie tematu). ### Firecrawl - niezawodne pobieranie stron Gdy wbudowany scraper (**scraper** - narzędzie, które pobiera treść strony) zawodzi, sięgam po `Firecrawl`. Zamienia dowolną stronę na czysty markdown albo strukturalny JSON, radząc sobie z JavaScriptem i podziałem na podstrony. ```bash pip install firecrawl-py ``` Ma pięć trybów: Scrape (jedna strona), Crawl (cała witryna po linkach), Map (lista adresów), Search (wyszukiwanie z treścią) i Extract (dane według schematu). Działa jako MCP (**MCP** - Model Context Protocol, standard podłączania narzędzi do agenta), więc wpina się prosto w Claude Code. ### Deep research - do przetestowania `NotebookLM-py` z sekcji o pamięci ma też tryb deep research po stronie Google, więc teoretycznie tańszy. Nie sprawdziłem go jeszcze na poważnie, dlatego traktuję to jako kandydata, nie rekomendację. ### Research swarm - struktura wpięta w wiki Do ustrukturyzowanego researchu używam skilla, który odpala agent swarm (**agent swarm** - rój niezależnych agentów pracujących równolegle). Jeden agent na pozycję, w partiach. Każdy zapisuje zwalidowany rekord JSON, a gotowy raport wpada do mojej bazy notatek. To zamienia jednorazowy research na trwały wpis w wiki. ## 🪙 Tokeny - gadatliwy output pali budżet i kontekst Rozwlekłe odpowiedzi (**output** - to, co agent wypisuje w odpowiedzi) kosztują podwójnie. Płacisz za tokeny i zapychasz okno kontekstowe. Im dłuższy output, tym szybciej agent traci miejsce na to, co ważne. ### Caveman - kompresja outputu `Caveman` to skill, który każe agentowi mówić jak jaskiniowiec: bez rodzajników, waty i uprzejmości, samo mięso. Tnie średnio **~65% tokenów outputu** (zakres 22-87%), zachowując 100% treści technicznej. ```powershell # Windows (PowerShell) irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex ``` ```bash # macOS / Linux / WSL curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash ``` Ma trzy poziomy: `/caveman lite`, `/caveman full` i `/caveman ultra`. I tu moje szczere zastrzeżenie: w trybie `full` opisy bywają nieczytelne, zwłaszcza po polsku. Dlatego trzymam się `lite`, a przy kodzie, commitach i tekstach o bezpieczeństwie przełączam na `normal mode`. Ważny szczegół z dokumentacji: Caveman ścina tylko output, nie tokeny myślenia (**reasoning** - tokeny, które model zużywa na wewnętrzne rozumowanie). Skraca usta, nie mózg. Największa korzyść to czytelność i szybkość, a oszczędność jest dodatkiem. ### Na radarze - kandydaci, nie rekomendacje Reszta zestawu leży u mnie na liście do sprawdzenia. Wymieniam je, ale wyraźnie zaznaczam, że sam ich jeszcze nie przetestowałem. - **Headroom** - kompresuje to, co agent czyta (input i kontekst) o 60-95%. Strona wejścia, tam gdzie Caveman zajmuje się wyjściem. - **Ponytail** - zmusza agenta do minimalnego kodu, więc mniej linii i mniej tokenów. - **Graphify** - zamienia kod w graf wiedzy, żeby agent nie czytał pliku po pliku. ## Kluczowe wnioski Pięć słabych stron, pięć konkretnych ruchów: 1. **Wideo** - dodaj `/watch`, zanim znów zaczniesz czytać sam transkrypt. Do generowania sprawdź HyperFrames. 2. **Design** - ustaw trzy skille w kolejności: kto (UX RULER), system (UI UX Pro Max), strażnik (Impeccable). 3. **Pamięć** - tekstowa wiki utrzymywana przez agenta bije RAG na małej skali. Jeśli budujesz od zera, zacznij od kursu. 4. **Research** - Firecrawl, gdy wbudowany scraper zawiedzie, a deep research tylko wtedy, gdy naprawdę go potrzebujesz. 5. **Tokeny** - Caveman ścina output, ale po polsku wybierz `lite`. Resztę traktuj jako kandydatów do testu. Najlepsze w tym wszystkim jest to, że każde z tych narzędzi jest darmowe i open source. Najgorsze, co się stanie, to że odpalisz jedno, nie spodoba Ci się i je odinstalujesz. Wybierz obszar, który najbardziej Cię boli, i zacznij od niego.

Chcesz wycisnąć więcej z Claude Code w swoim zespole?

Pomogę Ci dobrać narzędzia i skille pod Twój sposób pracy, wdrożyć je i ustawić powtarzalny proces pracy z agentem. A jeśli wolisz zacząć sam, pamięć agenta zbudujesz krok po kroku w darmowym kursie LLM Wiki.

Umów bezpłatną konsultację
## Przydatne zasoby - [Claude Video](https://github.com/bradautomates/claude-video) - skill `/watch`, który daje agentowi wzrok. - [HyperFrames](https://github.com/heygen-com/hyperframes) - generowanie wideo z plików HTML. - [Impeccable](https://impeccable.style/) - język designu plus linter usuwający AI slop. - [NotebookLM-py](https://github.com/teng-lin/notebooklm-py) - NotebookLM w terminalu Claude Code. - [Firecrawl](https://www.firecrawl.dev/) - dowolna strona zamieniona w czysty markdown. - [Caveman](https://github.com/JuliusBrussee/caveman) - kompresja tokenów w odpowiedziach agenta. - [Darmowy kurs LLM Wiki](/llm-wiki) - jak zbudować pamięć agenta (drugi mózg) od zera. ## FAQ
### W jakich obszarach Claude Code jest słaby out of the box? Claude Code świetnie radzi sobie z kodem, plikami i komendami, ale ma pięć słabych stron: nie widzi ani nie tworzy wideo, generuje generyczny design, zapomina po sesji, ma płytkie wyszukiwanie i gadatliwy output. Każdą z tych dziur łata się osobnym narzędziem albo skillem. To luki w funkcjach, a nie wady w pracy z kodem.
### Czy Claude Code potrafi obejrzeć wideo z YouTube? Domyślnie nie, ale zmienia to skill Claude Video z komendą `/watch`. Wkleja adres i pytanie, pobiera napisy oraz klatki wideo i czyta je jako obrazy, więc realnie „widzi" nagranie. Gdy brakuje napisów, zamienia mowę na tekst modelem whisper. Sprawdza się też do szybkiego pobierania transkryptów.
### Jak dać agentowi AI pamięć między sesjami? Najprostszy sposób to second brain, czyli tekstowa baza notatek w repozytorium, którą agent czyta, uzupełnia i przeszukuje. To wzorzec LLM Wiki: wiedza kumuluje się w plikach, inaczej niż w RAG, gdzie model odkrywa wszystko od nowa. Jak zbudować taką bazę od zera, pokazuję krok po kroku w [darmowym kursie LLM Wiki](/llm-wiki).
### Czym jest „AI slop" i jak uniknąć generycznego designu? AI slop to generyczny wygląd strony, który od razu zdradza, że front wygenerował model: ten sam układ hero, zaokrąglone karty i fioletowe gradienty. Unikniesz go, dając agentowi prawdziwy design system i linter, który wyłapie te wzorce. Używam do tego skilli UI UX Pro Max i Impeccable, ten drugi ma komendę `npx impeccable detect`.
### Czy Caveman naprawdę oszczędza tokeny i czy warto go używać? Tak, Caveman ścina średnio około 65% tokenów outputu, zachowując pełną treść techniczną. Zastrzeżenie: w trybie `full` opisy bywają nieczytelne, zwłaszcza po polsku, więc lepiej ustawić `lite`. Pamiętaj, że skraca tylko output, a nie tokeny myślenia modelu.
### Czy te narzędzia dla Claude Code są darmowe? Tak, wszystkie pięć rozwiązań jest darmowych i open source. Część wymaga własnych kluczy albo dodatków: Claude Video może potrzebować klucza whisper przy wideo bez napisów, a Firecrawl konta do trybu API. Instalacja sprowadza się zwykle do jednej komendy w terminalu.
--- # RAG ragowi nierówny: do kodu graf, do notatek indeks Source: https://pawel.lipowczan.pl/blog/rag-ragowi-nierowny Published: 2026-07-09 ## „RAG to głównie bazy wektorowe" - obiekcja, która ma rację Przy okazji [darmowego kursu LLM Wiki](/llm-wiki) dostałem od technicznego rozmówcy uwagę, którą parafrazuję: „RAG kojarzy mi się głównie z bazami wektorowymi. Ale do kodu są ciekawsze rozwiązania, oparte na analizie składniowej - na przykład codebase-memory-mcp, który autorzy nazywają Hybrid LSP. Do notatek pewnie średnio, ale do kodowania sprawdza się nieźle." Ma rację. W obu połowach zdania. Zamiast tę uwagę zbijać, wolę ją rozwinąć - siedzi w niej mapa, której brakuje większości dyskusji o **RAG** (Retrieval-Augmented Generation, czyli technice, w której model przed odpowiedzią dociąga treść z zewnętrznego źródła, zamiast polegać wyłącznie na pamięci z treningu). „Dodaj RAG do projektu" brzmi jak decyzja. Nie jest nią. To tak, jakby powiedzieć „użyj bazy danych" i nie sprecyzować, czy chodzi o relacyjną, dokumentową czy graf. Po tym artykule dobierzesz mechanizm wyszukiwania do materiału: inny do kodu, inny do notatek. Dowiesz się też, kiedy zwykły grep (wyszukiwanie w plikach po dokładnym ciągu znaków) wygrywa z całą resztą. ## RAG to rodzina technik, nie jedna Skojarzenie „RAG = baza wektorowa" wzięło się stąd, że przez dwa lata tak wyglądała większość wdrożeń. Baza wektorowa przechowuje **embeddings** (liczbowe reprezentacje tekstu, które pozwalają szukać po podobieństwie znaczenia, nie po dokładnym słowie). To jeden smak RAG, nie jego definicja. Przyjmijmy szerszą definicję: RAG to każde rozwiązanie, które tnie zużycie **tokenów** (jednostek, w których model rozlicza tekst) i dryf odpowiedzi przez dociąganie właściwej treści. W tej definicji mieszczą się cztery substraty: - **leksykalny** - grep i BM25 (klasyczny algorytm oceny trafności tekstu), szukanie po dokładnych słowach; - **semantyczny** - wektory i podobieństwo znaczeń; - **strukturalny** - grafy symboli, definicji i wywołań, zbudowane z samego kodu; - **kurowany indeks** - utrzymywany przez agenta spis treści bazy, jak w [LLM Wiki](/blog/llm-knowledge-base-brain-karpathy). Mój second brain (baza wiedzy w plikach markdown, utrzymywana przez agenta) też jest w tej definicji RAG-iem. Mechanizm ma tylko inny: zamiast liczyć podobieństwo wektorów, agent czyta lekki indeks i dociąga 2-3 trafne noty. Spór nie toczy się o cel. Toczy się o mechanizm - a **mechanizm ma pasować do materiału**. ## Kod parsuje się do symboli, notatki nie Kod różni się od notatek jedną cechą, która zmienia wszystko: jego gramatyka jest formalna i parser rozstrzyga ją jednoznacznie. Wywołanie `user.profile.name()` da się jednoznacznie rozwiązać - wiadomo, gdzie jest definicja, jaki typ wraca, kto jeszcze to woła. Z takiego materiału można zbudować **AST** (abstrakcyjne drzewo składniowe - strukturę, na którą parser rozkłada kod), a z wielu drzew **graf wywołań** (mapę „kto woła kogo" w całym repozytorium). Ta struktura jest dokładna i weryfikowalna. Notatki też mają gramatykę - języka naturalnego - ale maszyna nie rozstrzygnie jej jednoznacznie. Zdanie „decyzja była dobra, bo klient miał już PostgreSQL" nie rozkłada się na symbole: nie ma definicji, typów ani jednoznacznych referencji. Dlatego do notatek działa kurowany indeks plus linki między notami, a nie graf składniowy. Najostrzej ujęli to badacze z Duke i Snowflake w pracy HOMER o pamięci agentów: **similarity ≠ causality**. Bliskość dwóch fragmentów w przestrzeni embeddings to nie to samo, co zależność między nimi. Na pytanie „kto woła tę funkcję" wektor odpowiada zgadywaniem po podobieństwie, a graf faktami. Mój rozmówca zauważył to samo od strony praktyki: „do notatek średnio, do kodowania nieźle". Kod ma strukturę, którą da się zindeksować bez zgadywania - notatki nie. ## Studium: codebase-memory-mcp Narzędzie z obiekcji dobrze pokazuje, jak wygląda dojrzały retrieval do kodu. `codebase-memory-mcp` (DeusData) to serwer **MCP** (Model Context Protocol - standard podpinania narzędzi do agenta AI), który indeksuje całe repozytorium do trwałego grafu wiedzy o kodzie. Wbrew etykiecie „analiza składniowa" to hybryda trzech warstw: 1. **Graf symboli** - parser tree-sitter (rozkłada kod na AST, obsługuje 158 języków) plus „Hybrid LSP", czyli lekka reimplementacja tego, co robi **LSP** (Language Server Protocol - silnik, który w edytorze obsługuje „go to definition" i „find references"): rozwiązywanie typów, importów i dziedziczenia bez uruchamiania pełnego language servera. 2. **Wektory** - wkompilowane embeddings `nomic-embed-code`, bez klucza API. 3. **Pełny tekst** - wyszukiwanie BM25 po słowach. Sedno: narzędzie samo używa wektorów, tylko nie od nich zaczyna. Zaczyna od struktury, a podobieństwo semantyczne zostawia na rozmyte pytania. Liczba z nagłówka README robi wrażenie: **3 400 tokenów zamiast 412 000** na tym samym zadaniu. Pięć zapytań do grafu zamiast przeszukiwania repozytorium grep-em plik po pliku - redukcja o **99,2%**. Jedno zastrzeżenie: to deklaracja autorów (README plus preprint na arXiv), nie niezależny pomiar. Kierunek zgadza się jednak z tym, co znam z własnej bazy notatek: tam przejście z „przeszukaj wszystko" na „najpierw indeks" tnie koszt jednego pytania z ~35 tys. do ~3 tys. tokenów, około **30 razy**. Dwa różne materiały, dwie pary liczb, ten sam wzorzec: struktura przed siłowym czytaniem. ```text Pytanie: „co pęknie, gdy zmienię sygnaturę update_profile()?" grep plik po pliku: dziesiątki wyszukiwań i odczytów -> ~412 000 tokenów graf symboli: 5 zapytań strukturalnych -> ~3 400 tokenów (liczby: deklaracja autorów codebase-memory-mcp) ``` Uczciwie o granicach: pełne rozwiązywanie typów działa dla Pythona, TypeScript/JavaScript, Go, Rusta, Javy, C, C++, C#, Kotlina i PHP. Dla pozostałych języków narzędzie przechodzi na dopasowanie tekstowe - odpowiedź dostaniesz, ale mniej precyzyjną. ## Mapa substratów retrievalu codebase-memory-mcp to jeden punkt na większej mapie. Przegląd badań nad retrievalem do kodu (arXiv:2510.04905) dzieli metody na **graph-based** (rdzeniem jest jawnie zbudowany graf struktury kodu) i **non-graph** (repozytorium traktowane jak worek tekstu, przeszukiwany leksykalnie albo semantycznie). Na tej osi układa się pięć praktycznych podejść: | Substrat | Jak działa | Przykłady | |----------|-----------|-----------| | Graf AST (tree-sitter) | parsuj kod, zbuduj graf definicji i wywołań | codebase-memory-mcp, Graphify, repo map Aidera | | LSP jako kontekst | agent pyta language server o symbole i typy | Serena, mcp-language-server | | Utrwalone fakty (SCIP/LSIF) | fakty o kodzie zapisane i odpytywane jak baza danych | Glean (Meta), srctx | | Agentic grep | agent sam generuje i wykonuje wyszukiwania, zero indeksu | Claude Code, Cline, GrepRAG | | Wektory (semantic) | podobieństwo embeddings | klasyczny RAG | | Kurowany indeks przy kodzie | drzewo plików `AGENTS.md`, agent schodzi od korzenia do miejsca edycji | DOX | **SCIP i LSIF** to formaty utrwalania faktów o kodzie (kto definiuje, kto używa), zapisywalne w repozytorium i współdzielone w zespole. Szósty wiersz wymaga słowa wyjaśnienia: `AGENTS.md` to plik instrukcji, który agent kodujący czyta automatycznie na starcie pracy w repozytorium. Wzorzec DOX (repo `agent0ai/dox`) mnoży go - jeden `AGENTS.md` w każdym folderze, z przeznaczeniem folderu, lokalnymi konwencjami i indeksem podfolderów - a agent schodzi tym drzewem od korzenia do miejsca edycji i po zmianie aktualizuje dokumentację. To ten sam kurowany indeks, co przy notatkach, tylko wersjonowany razem z kodem. Podejść jest więcej - grafy ścieżek wykonania, hybrydy z wiedzą zespołu - ale ta szóstka pokrywa większość decyzji, które realnie podejmiesz. ## Uczciwie: kiedy grep bije graf Gdyby struktura zawsze wygrywała, ten artykuł byłby krótszy. Nie wygrywa. Preprint GrepRAG (Uniwersytet Zhejiang, styczeń 2026) odwraca perspektywę. Zamiast budować jakikolwiek indeks, model sam generuje około 10 komend `ripgrep`, wykonuje je na surowym repozytorium i przestawia wyniki wagami BM25. Zero indeksu oznacza zero nieświeżego indeksu - nic nie rozjeżdża się z kodem. Według autorów podejście dorównało grafom i wektorom albo je pobiło na uzupełnianiu kodu, przy retrievalu średnio ~13 razy szybszym (w skrajnym przypadku 35 razy). To preprint z wynikami własnymi autorów - traktuję go jak deklarację, nie fakt. Mocniejszy argument przyszedł z praktyki. Zespół Claude Code porzucił wczesną wersję opartą na RAG z lokalną bazą wektorową, bo wyszukiwanie agentowe - grep plus nawigacja po plikach - działało lepiej i nie miało problemu przeterminowanego indeksu. Cline, inny popularny agent kodujący, ogłosił to samo stanowisko w manifeście „why we don't index your codebase". Po drugiej stronie stoi RepoGraph (ICLR 2025, praca recenzowana): dołożenie grafu repozytorium poprawiło wyniki czterech agentów średnio o **32,8%** na SWE-bench-Lite (benchmarku naprawiania prawdziwych błędów z projektów GitHub). Grep kontra graf - kto ma rację? Oba obozy. Rozstrzyga zadanie: | Zadanie | Wygrywa | Dlaczego | |---------|---------|----------| | Lokalne uzupełnienie fragmentu | agentic grep | tanio, bez indeksu, bez dryfu - odpowiedź jest blisko | | Zmiana przechodząca przez wiele plików | graf / LSP | potrzebna mapa wywołań i typów, której lokalne okno nie pokaże | | „Kto woła X? Co pęknie po zmianie?" | LSP / graf definicji | wynik wyczerpujący, nie ranking kandydatów | | Rozmyte „gdzie jest logika logowania" | wektory | dopasowanie intencji, nie dokładnej struktury | W praktyce nie wybierasz więc „graf albo grep". Grep bierze lokalną robotę, struktura zmiany przekraczające granice modułów, wektory rozmyte pytania o intencję. codebase-memory-mcp po cichu przyznaje to samo: w jednym narzędziu siedzi graf, embeddings i grep wzbogacony o kontekst z grafu. Jedną lukę muszę odnotować. Wszystkie te narzędzia są zweryfikowane na Pythonie, TypeScript, Go czy Ruście. Na C# i Roslyn (kompilatorze .NET) nie znalazłem potwierdzonego wdrożenia żadnego z nich. Mosty LSP powinny zadziałać w teorii; w praktyce to test do zrobienia, nie fakt do zacytowania. Sam już prawie nie piszę kodu ręcznie - robią to agenci - ale to nie unieważnia problemu: retrieval, którym agent czyta repozytorium, wciąż zależy od wsparcia konkretnego języka. Jeśli Twój zespół siedzi w .NET, przetestuj, zanim zaufasz benchmarkom. ## Dwie warstwy pamięci agenta Wróćmy do obiekcji. „Do kodu structural, do notatek średnio" - zgoda. Wniosek nie brzmi jednak „wybierz jedno". W agencie kodującym chcesz obu warstw naraz, bo pamiętają co innego. **Graf i LSP pamiętają, jak wygląda kod**: co się z czym woła, gdzie jest definicja, co pęknie po zmianie sygnatury. **Second brain pamięta, dlaczego kod tak wygląda**: jaką konwencję przyjęliśmy, czego próbowaliśmy wcześniej i co nie zadziałało, dlaczego wybraliśmy tę bibliotekę, a nie tamtą. Graf da Ci pełne drzewo wywołań. Nie powie, że obecną konwencję narzuciliśmy po tym, jak poprzednia wysadziła produkcję w marcu. Tej informacji nie ma w żadnym AST, bo nikt jej tam nie zapisał. Została w głowach, w wątkach na Slacku albo - jeśli prowadzisz bazę wiedzy - w jednej nocie z datą i uzasadnieniem. Kod nie zapisuje intencji. Brain tak. Gdzie w tym podziale mieszka drzewo `AGENTS.md`? Pomiędzy. Lokalne konwencje folderu - „tu testy piszemy tak, tego modułu nie ruszaj" - najlepiej trzymać tuż przy kodzie, wersjonowane razem z nim, i dokładnie to robi wzorzec DOX. Brain trzyma piętro wyżej to, czego pojedyncze repozytorium nie pomieści: decyzje i standardy wspólne dla wielu projektów, historię podejść, wnioski z porażek. Z dwóch warstw robią się trzy: graf pamięta strukturę, `AGENTS.md` lokalne konwencje, brain wiedzę ponad projektem. Dlatego mój rozmówca miał rację w obu połowach zdania, a jego obiekcja opisuje nie konkurencję, tylko podział pracy między dwie warstwy pamięci jednego agenta. Strukturę kodu trzymaj w grafie. Decyzje, konwencje i wnioski w bazie notatek - [przenośnej i czytelnej dla agenta](/blog/okf-standard-przenosnosc-bazy-wiedzy-ai). Jak taką bazę zbudować od zera, pokazuję w [darmowym kursie LLM Wiki](/llm-wiki). ## Kluczowe wnioski 1. **RAG ragowi nierówny.** Wybieraj substrat: leksykalny (grep/BM25), semantyczny (wektory), strukturalny (AST/LSP/graf) albo kurowany indeks. Baza wektorowa to jeden smak, nie definicja. 2. **Kod parsuje się do symboli, notatki nie.** Do zmian przekraczających granice plików domyślnie używaj struktury; do notatek indeksu i linków. 3. **Grep to zawodnik, nie rezerwa.** Brak indeksu bije nieświeży indeks przy lokalnej robocie - Claude Code i Cline zbudowały na tym całe podejście. 4. **Dopasuj mechanizm do zadania, nie do mody.** Benchmarki z README i preprintów traktuj jak deklaracje autorów, dopóki ktoś ich niezależnie nie powtórzy. 5. **Agentowi kodującemu daj obie warstwy pamięci.** Graf pamięta strukturę, brain pamięta intencje, a lokalne konwencje folderu trzymaj przy kodzie (drzewo `AGENTS.md`). Dopiero razem odpowiadają i na „co pęknie", i na „dlaczego tak".

Zastanawiasz się, jakim retrievalem karmić swojego agenta AI?

Pomagam dobrać warstwy pamięci agentów do realnego projektu: graf do kodu, bazę wiedzy do decyzji i konwencji. Powiem Ci, co zadziała w Twoim zespole, zanim kupisz kolejne narzędzie.

Umów bezpłatną konsultację

Warstwę wiedzy zbudujesz sam - w darmowym kursie LLM Wiki pokazuję jak, krok po kroku. Szablon z kursu połączysz z repozytorium kodu.

## Przydatne zasoby - [codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp) - graf kodu jako serwer MCP; preprint: [arXiv:2603.27277](https://arxiv.org/abs/2603.27277) - [Retrieval-Augmented Code Generation - przegląd metod](https://arxiv.org/abs/2510.04905) - taksonomia graph vs non-graph - [GrepRAG](https://arxiv.org/html/2601.23254v2) - retrieval bez indeksu: model sam generuje komendy ripgrep - [RepoGraph](https://arxiv.org/abs/2410.14684) - recenzowany dowód, że graf repozytorium poprawia wyniki agentów - [Cline: why we don't index your codebase](https://cline.bot/blog/why-cline-doesnt-index-your-codebase-and-why-thats-a-good-thing) - manifest podejścia agentic grep - [Serena](https://github.com/oraios/serena) - LSP jako narzędzia agenta (symbole, referencje, refaktoryzacja) - [DOX](https://github.com/agent0ai/dox) - samodokumentujące się drzewo `AGENTS.md`: kurowany indeks wersjonowany z kodem - [Repo map Aidera](https://aider.chat/docs/repomap.html) - protoplasta map repozytorium z budżetem tokenów - [Jak LLM Wiki Karpathy'ego pomogła mi uporządkować bazę wiedzy](/blog/llm-knowledge-base-brain-karpathy) - kurowany indeks w praktyce - [Zbudowałem drugi mózg. Trafiłem w standard Google (OKF)](/blog/okf-standard-przenosnosc-bazy-wiedzy-ai) - przenośność bazy notatek ## FAQ
### Czy RAG to zawsze baza wektorowa i embeddings? Nie. Baza wektorowa to jeden z czterech substratów retrievalu, obok leksykalnego (grep, BM25), strukturalnego (grafy AST/LSP) i kurowanego indeksu. Jeśli definiować RAG jako dociąganie właściwej treści przed odpowiedzią modelu, to LLM Wiki z indeksem też jest RAG-iem - tylko bez wektorów.
### Kiedy do przeszukiwania kodu wystarczy grep, a kiedy potrzebny jest graf symboli? Grep wystarcza przy lokalnej robocie: uzupełnienie fragmentu, znany ciąg znaków, odpowiedź blisko miejsca edycji. Graf albo LSP potrzebujesz przy zmianach przekraczających granice plików i pytaniach „kto woła X, co pęknie po zmianie". Tam grep daje fałszywe trafienia, a wektory zgadują.
### Czym różni się retrieval do kodu od retrievalu do notatek? Kod ma gramatykę formalną, którą parser rozstrzyga jednoznacznie, więc można z niego zbudować dokładny graf definicji, typów i wywołań. Tekst notatki nie rozkłada się na symbole - zdanie „dlaczego podjęliśmy tę decyzję" nie ma definicji ani typów. Dlatego do notatek działa kurowany indeks i linki między notami, a do kodu struktura.
### Czy LLM Wiki zastępuje narzędzia typu codebase-memory-mcp? Nie, to komplementarne warstwy pamięci agenta. Graf kodu pamięta strukturę: wywołania, typy, zależności między plikami. LLM Wiki pamięta intencje: decyzje, konwencje, odrzucone podejścia. Agent kodujący korzysta z obu naraz.
### Co znaczy „similarity ≠ causality" przy bazach wektorowych? Bliskość dwóch fragmentów w przestrzeni embeddings nie oznacza, że jeden zależy od drugiego. Na pytania strukturalne, jak „kto woła tę funkcję", wektor odpowiada podobieństwem, czyli zgadywaniem, a graf wywołań faktami. Sformułowanie pochodzi z pracy HOMER (Duke + Snowflake) o pamięci agentów.
### Gdzie trzymać konwencje projektu - w plikach AGENTS.md czy w second brain? Lokalne konwencje folderu (styl testów, moduły, których nie wolno ruszać) trzymaj przy kodzie - w drzewie `AGENTS.md` wersjonowanym razem z plikami, jak we wzorcu DOX. W second brain trzymaj to, co przekracza jedno repozytorium: decyzje architektoniczne z uzasadnieniem, standardy wspólne dla projektów, historię odrzuconych podejść. Graf symboli uzupełnia oba - pamięta strukturę, której nie opisuje żaden markdown.
--- # Zbudowałem drugi mózg. Trafiłem w standard Google (OKF) Source: https://pawel.lipowczan.pl/blog/okf-standard-przenosnosc-bazy-wiedzy-ai Published: 2026-06-21 ## Zbudowałem drugi mózg. Okazało się, że w 100% trafiłem w standard Google (OKF) Jakiś czas temu opisałem, [jak LLM Wiki Karpathy'ego pomogła mi uporządkować moją bazę wiedzy](/blog/llm-knowledge-base-brain-karpathy) - ponad 300 notatek, którymi zarządza agent, trzy indeksy nawigacyjne, workflow do ingestu i kompilacji. System działa. Codziennie. To mój drugi mózg. Ale ostatnio złapałem się na niewygodnym pytaniu. Cała ta wiedza żyje w **moich** narzędziach - w Obsidian, w Quartz, w moim `CLAUDE.md`. A co, jeśli zechcę ją komuś przekazać? Zmienić tool? Albo potraktować jako aktyw firmowy, który ma przeżyć dowolny program i dowolny model? Notatki zamknięte w czyimś prywatnym formacie to vendor-lock-in - tyle że zrobiony własnoręcznie. Wniosek, do którego doszedłem, jest prosty: **trwałość wiedzy to nie narzędzie - to format.** I wtedy trafiłem na **OKF (Open Knowledge Format)** od Google. Zrobiłem audyt swojego brain pod kątem tego standardu, spodziewając się sporego rozjazdu. Wynik mnie zaskoczył: zgodność w okolicach 100% - mimo że nigdy nie projektowałem brain pod OKF. W tym artykule pokażę, dlaczego tak się stało, co dokładnie sprawdziłem i co wtedy, gdy czysty markdown to za mało. Bez teorii - konkretny audyt i konkretne liczby. ## Czym jest OKF (Open Knowledge Format) `OKF` to otwarty format zapisu wiedzy, zaprojektowany pod erę agentów. Technicznie nie ma w nim magii: to **markdown + frontmatter YAML**. Cała filozofia mieści się w jednym haśle wzorca - _„authored by people, generated by agents"_ (tworzą ludzie, generują i utrzymują agenci). Dokładnie ten podział, który u siebie wypracowałem organicznie. Podstawową jednostką jest `Knowledge Bundle` - katalog plików `.md`. Może być repozytorium git, tarballem albo podkatalogiem. I to jest sedno przenośności: `git clone` i masz całość. Wewnątrz `Concept` to jedna myśl - **jeden plik `.md` = jeden koncept** (frontmatter plus treść markdown). Format rezerwuje dwie nazwy plików: `index.md` (listing katalogu pod [progressive disclosure](/blog/llm-knowledge-base-brain-karpathy)) oraz `log.md` (chronologia zmian, najnowsze u góry). Najciekawsza jest minimalna bariera wejścia. **Jedyne twardo wymagane pole frontmattera to `type`.** Cała reszta - `title`, `description`, `resource`, `tags`, `timestamp` - jest zalecana, ale opcjonalna. Producent może dodawać własne klucze, a konsument **musi** tolerować nieznane pola i broken links. Najprostszy zgodny koncept wygląda tak: ```yaml --- type: note --- Treść notatki w czystym markdownie. ``` Warto rozumieć różnicę między OKF a `LLM Wiki`. **OKF to specyfikacja interoperacyjności** - mówi „jak zapisać wiedzę, żeby była wymienialna". **LLM Wiki to metodologia pracy** - mówi „jak agent ma tę wiedzę budować i utrzymywać". To dwie komplementarne warstwy, a nie konkurencja. Jedna uczciwa uwaga. Repozytorium OKF ma wprost notkę: _„This repository and its contents are not an official Google product"_. OKF to otwarta specyfikacja interoperacyjności, a nie produktowe zobowiązanie Google. ## Mój brain trafił w standard - choć go pod niego nie projektowałem To jest serce tej historii. Zrobiłem realny audyt zgodności brain z OKF v0.1. Werdykt jednym zdaniem: **brain przechodzi 100% twardych warunków conformance OKF v0.1; rozbieżności są wyłącznie w warstwie rekomendowanej, czyli kosmetyce.** Spec definiuje trzy twarde warunki zgodności. Brain spełnia wszystkie trzy: | # | Wymóg OKF | Status | Dowód w brain | | --- | ------------------------------------------------------------ | :----: | --------------------------------------------------------------------------------------------------------------- | | 1 | Każdy nie-zarezerwowany `.md` ma parsowalny YAML frontmatter | ✅ | Każda nota otwiera się `--- ... ---` | | 2 | Każdy frontmatter ma niepuste `type` | ✅ | `type: basic-note \| book-note \| knowledge-note \| tool \| compiled-note \| answer-note` | | 3 | Pliki zarezerwowane mają właściwą strukturę _gdy istnieją_ | ✅ | Brain nie używa literalnych `index.md`/`log.md` → funkcję pełnią `_indexes/*`; warunek „gdy istnieją" spełniony | Liczby na głos: **twarda zgodność = 100% (3/3)**. Warstwa rekomendowana (te wszystkie opcjonalne pola i konwencje) to **~85-90%**. I tu robi się ciekawie, bo wszystkie cztery rozbieżności są nazewniczo-składniowe - dotyczą tego, _jak coś nazwałem_, a nie tego, _jak modeluję wiedzę_: | Obszar | OKF | brain | Charakter | | ------------------- | ---------------------------------------- | --------------------------------------------------- | ------------------------------------------------------------ | | Linki | `[label](/path)` | `[[wikilinks]]` (Obsidian) | składnia; **Quartz kompiluje do `[label](path)` w buildzie** | | Indeks | literalny `index.md` | `_indexes/vault-map.md` + `catalog.md` + `graph.md` | nazwa, nie funkcja | | Log | `log.md` | sekcja `## Recent Changes` w `_indexes/vault-map.md` + historia git | rola w sekcji, nie w osobnym pliku | | Klucze frontmattera | `description` / `resource` / `timestamp` | `summary` / `source` / `date` | mapowalne 1:1 przy eksporcie | Spójrz na kolumnę „Charakter". Nigdzie nie ma „brakuje mi tej koncepcji" - wszędzie jest „nazwałem to inaczej" albo „narzędzie kompiluje to za mnie". Wikilinks Obsidian są pełnoprawnymi linkami OKF po buildzie Quartz. `_indexes/` pełni dokładnie rolę `index.md`. Pola mapują się jeden do jednego. > „Zgodność nie jest przypadkiem - brain i OKF wyrastają z tej samej koncepcji LLM Wiki Karpathy'ego. OKF to spec interoperacyjności, brain to działająca implementacja. Różnice to wybory narzędzia, nie różnice modelu wiedzy." To jest **konwergentna ewolucja**. Dwa systemy budowane niezależnie - specyfikacja w Google i mój vault w Obsidian - zbiegły się do tego samego kształtu, bo wyrosły z tej samej idei. I właśnie dlatego standard jest wiarygodny: nie wymyślono go zza biurka, tylko spisano to, do czego praktyka i tak dąży. ## Dlaczego format ma znaczenie Można zapytać: skoro mój system działa, po co mi w ogóle standard? Odpowiedź to pięć konkretnych rzeczy, które daje format - a których nie da żadne pojedyncze narzędzie: - **Przenośność.** _„If you can `git clone` it, you can ship it."_ Bundle to repozytorium - kopiujesz całość i działa u kogokolwiek, bez instalacji, bez konfiguracji, bez zależności od mojego setupu. - **Interoperacyjność.** Bundle wiedzy można wymieniać niezależnie od narzędzia. Eksport z mojego brainu, import do cudzego. To realny przyszły kierunek produktowy - ekstrakt bazy jako gotowy bundle OKF. - **Trwałość i future-proofing.** Narzędzia umierają, format zostaje. Markdown plus frontmatter przeżyje Obsidian, przeżyje Quartz, przeżyje konkretny model LLM. Za pięć lat wciąż otworzysz te pliki. - **Czytelność podwójna.** Ten sam plik czyta agent i człowiek. Bez żadnego toola wystarczy go otworzyć w edytorze - to nie baza w zamkniętym formacie binarnym. - **Aktyw, nie silos.** Wiedza zapisana w standardzie staje się zasobem, który da się audytować, przekazać i wycenić. I to jest most do kolejnego pytania - co, gdy ten aktyw rośnie do skali firmy. ## Gdy czysty markdown to za mało - Google Knowledge Catalog Tu zmieniam ton na ostrożny. Tego fragmentu **nie testowałem** - sygnalizuję kierunek, nie wystawiam rekomendacji. Karpathy sam zakreśla skalę, na której to podejście działa: _„moderate scale (~100 sources, ~hundreds of pages)"_ - zanim potrzebny staje się embedding-based search. Mój brain to dziś rzędu 300 notatek, czyli wciąż ten przedział: czysty markdown z indeksami ogarnia go bez embeddings i bez infrastruktury RAG. W [poprzednim artykule](/blog/llm-knowledge-base-brain-karpathy) pokazywałem ten mechanizm w liczbach: progressive disclosure daje tam około **30× mniej kontekstu** na zapytanie niż wrzucanie całego vaulta. To nie twardy limit - raczej rząd wielkości, do którego index-first pozostaje wygodny. Na osobisty drugi mózg w zupełności wystarcza. Ale powyżej - przy milionach dokumentów, danych ustrukturyzowanych i nieustrukturyzowanych jednocześnie, wielu agentach pracujących równolegle - to inna liga. Tu wchodzi **Google Knowledge Catalog** (managed, wg strony produktu „formerly Dataplex"). Google opisuje go jako _„universal context engine for your enterprise"_ - „always-on context and governance for your agents". Zamiast Twoich notatek kataloguje cały data estate: automatycznie zbiera metadane z BigQuery, AlloyDB, Spannera czy Lookera, a źródła nieustrukturyzowane - PDF-y, kontrakty, wiki - zamienia w _„structured knowledge graph"_, który agent może odpytać. To samo repozytorium (`GoogleCloudPlatform/knowledge-catalog`), które zawiera katalog `okf/`, trzyma też `samples/` i `toolbox/`. Widać tu ciągłość, nie przeskok. Ta sama idea - wiedza ustandaryzowana, semantyczna, czytelna dla agentów - tyle że na poziomie przedsiębiorstwa. OKF i LLM Wiki to skala osobista; Knowledge Catalog to skala firmowa. Jeden łańcuch, dwa końce. **Gdzie dokładnie jest granica?** To nie jest „local vs cloud" - mój brain też stoi w chmurze. Różnica jest jakościowa, na czterech osiach: - **Co katalogujesz.** Brain = Twoje autorskie notatki. Catalog = warstwa semantyczna nad całym, żywym data estate firmy: tabele, hurtownie, pliki, modele BI. - **Silnik retrieval.** Brain = index-first plus progressive disclosure, bez embeddings. Catalog = semantic search z sub-sekundową latencją nad knowledge graphem milionów encji. - **Governance.** Brain = git i konwencje. Catalog = polityki dostępu (IAM), data quality, lineage, audyt - retrieval respektuje uprawnienia, więc agent widzi tylko to, do czego ma prawo. - **Świeżość i współbieżność.** Brain = snapshot, który sam utrzymujesz. Catalog = always-on, aktualizuje się wraz z danymi i obsługuje wielu agentów naraz przez Context API i MCP tools. Liczba plików to tylko objaw. Prawdziwa granica to **typ danych plus silnik retrieval plus governance plus dynamika**. Dla osobistej wiedzy w prozie czysty markdown wygrywa prostotą; dla firmowego, heterogenicznego data estate potrzebujesz warstwy pokroju Knowledge Catalog. I jeszcze raz, bo to ważne: repozytorium OKF nosi notkę _„not an official Google product"_, a Knowledge Catalog opisuję tu wyłącznie jako kierunek, którego sam nie wdrażałem. Bez obietnic co do funkcji produktu. ## Co z tego wynika 1. **Format wygrywa z narzędziem.** Jeśli budujesz bazę wiedzy, projektuj ją pod standard od pierwszego dnia. Migracja później kosztuje - zgodność od startu jest darmowa. 2. **Zgodność bywa konwergentna.** Jeśli twój system wyrasta z dobrej idei (LLM Wiki), masz szansę trafić w standard bez celowania w niego. To dobra wiadomość: nie musisz przepisywać brainu, prawdopodobnie wystarczy go wyeksportować. 3. **Wiedza w standardzie to aktyw.** Przenośny, audytowalny, wymienialny - i skalujący się od osobistego brain po Knowledge Catalog. Jeśli chcesz zacząć z gotowym fundamentem, udostępniam **[szablon `second-brain-template`](https://github.com/plipowczan/second-brain-template)** - „Use this template", `/onboard`, pierwszy ingest i masz własny system zgodny ze standardem w kilka minut. Pracuję też nad szerszymi materiałami o budowaniu drugiego mózgu - jeśli temat cię interesuje, warto trzymać rękę na pulsie.

Chcesz bazę wiedzy, która jest Twoja na zawsze - przenośna i gotowa dla agentów?

Pomogę Ci zaprojektować architekturę bazy wiedzy zgodną ze standardem (OKF) - od struktury notatek i indeksów po skalę firmową. Możesz też wziąć darmowy szablon i odpalić własny system w kilka minut.

Umów bezpłatną konsultację
## Przydatne zasoby - [Jak LLM Wiki Karpathy'ego pomogła mi uporządkować bazę wiedzy](/blog/llm-knowledge-base-brain-karpathy) - bazowy artykuł o LLM Wiki, progressive disclosure i architekturze brainu - [Second brain w Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills) - jak zacząć od zera - [brain.lipowczan.pl](https://brain.lipowczan.pl) - żywa instancja mojego drugiego mózgu - [second-brain-template](https://github.com/plipowczan/second-brain-template) - gotowy szablon do startu - [OKF / knowledge-catalog (repo)](https://github.com/GoogleCloudPlatform/knowledge-catalog) - specyfikacja OKF w katalogu `okf/` (uwaga: „not an official Google product") - [Google Knowledge Catalog](https://cloud.google.com/products/knowledge-catalog) - managed platforma katalogu danych (skala enterprise) ## FAQ
### Czym jest Open Knowledge Format (OKF) i kto za nim stoi? OKF (Open Knowledge Format) to otwarta specyfikacja zapisu wiedzy oparta na markdown plus frontmatter YAML, zaprojektowana pod pracę z agentami AI. Specyfikacja żyje w repozytorium `GoogleCloudPlatform/knowledge-catalog` w katalogu `okf/`. Ważna uwaga: repozytorium nosi notkę „This repository and its contents are not an official Google product" - to otwarty standard interoperacyjności, nie produktowe zobowiązanie Google.
### Czym OKF różni się od LLM Wiki Karpathy'ego? OKF to specyfikacja interoperacyjności - definiuje, jak zapisać wiedzę, żeby była przenośna i wymienialna między narzędziami. LLM Wiki to metodologia pracy - opisuje, jak agent ma budować i utrzymywać bazę wiedzy. Są komplementarne: OKF mówi o formacie, LLM Wiki o procesie. Działająca baza wiedzy korzysta z obu warstw naraz.
### Jedyne wymagane pole w OKF to `type` - co to oznacza w praktyce? Oznacza minimalną barierę wejścia: zgodny plik OKF musi mieć tylko niepuste pole `type` we frontmatterze, a cała reszta (`title`, `description`, `tags`, `timestamp`) jest opcjonalna. Producent może dodawać dowolne własne klucze, bo konsument formatu musi tolerować nieznane pola i broken links. Dzięki temu standard jest liberalny przy zapisie, a rygorystyczny tylko tam, gdzie to konieczne.
### Czy moja istniejąca baza w Obsidian jest zgodna z OKF? Najprawdopodobniej w dużej części tak, zwłaszcza jeśli używasz plików `.md` z frontmatterem. W moim audycie brain przeszedł 100% twardych warunków conformance OKF v0.1 (3/3), a rozbieżności były wyłącznie kosmetyczne. Wikilinks `[[...]]` Obsidian kompilują się przez Quartz do linków `[label](path)`, a pola jak `summary` czy `date` mapują się jeden do jednego na `description` i `timestamp` przy eksporcie.
### Kiedy czysty markdown przestaje wystarczać i wchodzi Google Knowledge Catalog? Podejście index-first na czystym markdownie działa świetnie w skali osobistej - rzędu setek notatek (mój brain to dziś nieco ponad 300), bez embeddings i bez infrastruktury RAG. To nie twardy limit, tylko rząd wielkości. Powyżej - przy milionach dokumentów, danych ustrukturyzowanych i nieustrukturyzowanych oraz wielu agentach naraz - wchodzi skala enterprise, którą adresuje Google Knowledge Catalog (managed, „formerly Dataplex"). Zaznaczam jednak, że tego progu osobiście nie testowałem - sygnalizuję kierunek, nie rekomendację produktową.
--- # Software 3.0: dlaczego twoja aplikacja nie powinna istnieć Source: https://pawel.lipowczan.pl/blog/software-3-0-agentic-engineering Published: 2026-05-29 Człowiek, który ukuł termin **vibe coding**, powiedział, że „nigdy nie czuł się bardziej w tyle jako programista". To Andrej Karpathy - współzałożyciel OpenAI, były szef AI w Tesli. W grudniu agentowe modele przekroczyły dla niego pewien próg: kawałki kodu „po prostu wychodziły dobrze", więc przestał je poprawiać i zaufał systemowi. To nie jest historia o tym, że kod pisze się szybciej. To historia o tym, że **programowanie zmieniło się w prompcie** - a jeśli czytasz „software" jako wszystkie cyfrowe biznesy, to znaczy, że twój produkt właśnie jest przebudowywany na AI-first, czy tego chcesz, czy nie. Przez ostatnie miesiące buduję z agentami codziennie - całe portfolio, na którym czytasz ten tekst, i systemy w dwóch firmach. Im więcej z nimi pracuję, tym mocniej widzę, że to nie jest „kolejny hype o AI". To zmiana mapy: rośnie abstrakcja, a wraz z nią przesuwa się to, co faktycznie robi człowiek. Ten artykuł to próba narysowania tej mapy - żebyś wiedział, w którym paradygmacie działasz, gdzie jest twoja fosa i czego nie da się oddelegować. ![Diagram The Karpathy Paradigm - pięć pasów: ① trzy paradygmaty Software 1.0→2.0→3.0 (rosnąca abstrakcja), ② vibe coding podnosi podłogę vs agentic engineering trzyma sufit, ③ weryfikowalność i jagged intelligence (refactor 100k-line vs car wash), ④ co zostaje po stronie człowieka - taste, judgment, understanding, ⑤ buduj dla agentów: sensory, aktuatory, dane legible dla LLM](/images/karpathy-paradigm-software-3-0.webp) Ten diagram to oś całego tekstu - wracam do każdego z pięciu pasów w kolejnych sekcjach. ## Trzy paradygmaty: jak „programujesz" maszynę Karpathy układa ewolucję oprogramowania w trzy paradygmaty - i kluczem jest to, *jak* przekazujesz maszynie swoją intencję. - **Software 1.0** - piszesz jawny kod, regułę po regule. Krok pierwszy, krok drugi, krok trzeci. Kruche, nie samonaprawialne: gdy coś wypadnie poza scenariusz, skrypt się wywraca. - **Software 2.0** - „programujesz" przez **dane**. Nie piszesz reguł, tylko kurujesz zbiory i trenujesz wagi sieci neuronowej. Logika siedzi w wagach, nie w `if`-ach. - **Software 3.0** - **prompting**. Okno kontekstu (context window) to twoja dźwignia nad interpreterem, którym jest LLM. > „Software 3.0 is kind of about your programming now turns to prompting. And what's in the context window is your lever over the interpreter that is the LLM." - Andrej Karpathy *(Programowanie zmienia się w pisanie promptów, a to, co masz w oknie kontekstu, jest twoją dźwignią nad LLM-em jako interpreterem.)* I tu pada reframe, który zmienia ten tekst z „ciekawostki dla programistów" w coś, co dotyczy każdego, kto buduje produkt. Czytaj „software" jako **wszystkie cyfrowe biznesy**. LLM to silnik - a ty projektujesz samochód wokół niego. > „All businesses are literally being restructured to be AI first... AI in the middle being the engine and driving them forward. But that is all the engine is. You get to design your own car." *(Wszystkie biznesy są przebudowywane na AI-first - z AI jako silnikiem w środku. Ale to tylko silnik. To ty projektujesz swój samochód.)* To nie jest abstrakcja. To samo widać w tym, jak firmy zaczynają działać jako systemy agentowe - z kontekstem jako trwałym aktywem i software'em jako warstwą, którą się regeneruje, gdy pojawia się lepszy model. ## Cztery przykłady, które to unaoczniają Paradygmat brzmi abstrakcyjnie, dopóki nie zobaczysz, co robi z konkretnymi modelami biznesowymi. Karpathy daje cztery przykłady i każdy z nich to inny typ produktu, który właśnie się przesuwa. 1. **Instalator.** Zamiast pęczniejącego skryptu („krok pierwszy: zaakceptuj warunki, krok drogi: znajdź folder..."), wywołujesz agenta jedną komendą. On *posiada cel* („zainstaluj to") i sam debuguje w pętli problemy, których nigdy wcześniej nie widział. To mały, samonaprawialny „skill" - blob tekstu, który wklejasz swojemu agentowi. 2. **Aplikacja (MenuGen).** Karpathy zbudował apkę 2.0: robisz zdjęcie menu, a ona renderuje obrazy dań. Stała się „natychmiast bezużyteczna" - bo wystarczy podać zdjęcie do Gemini i powiedzieć „nałóż obrazy pozycji przez Nano Banana". Nie ma czego instalować. *„Ta aplikacja nie powinna istnieć."* I to samo nadchodzi po niemal każdą pojedynczą apkę. 3. **Kurs / wiedza ekspercka.** Od sekwencyjnych lekcji na Udemy, przez kurs przemodelowany danymi o zaangażowaniu, do agenta-coacha, który *robi to z tobą*, kiedy działasz. Nie „zadaj mi pytanie", tylko „zbuduj to ze mną". 4. **Usługa (montaż wideo).** Od ręcznego Premiere Pro, przez Descript wycinający ciszę, do pola tekstowego: „zmontuj to w wiralowym stylu MrBeast, poniżej 8 minut" - i gotowy render w kilka minut. Wspólny mianownik jest brutalnie prosty: **Software 3.0 sprzedaje wynik (outcome), nie narzędzie.** Jeśli twój produkt jest narzędziem, które wykonuje krok, a nie dostarcza efektu - to jest na celowniku. ## Vibe coding podnosi podłogę. Agentic engineering trzyma sufit. To jest moim zdaniem najważniejsze nowe rozróżnienie z całego wywiadu - i miejsce, gdzie najwięcej osób się myli. **Vibe coding** i **agentic engineering** to dwie różne rzeczy. **Vibe coding podnosi PODŁOGĘ.** Każdy może teraz zbudować software. To demokratyzujące i naprawdę świetne - opisałem zresztą całe podejście w [przewodniku po vibe codingu](/blog/vibe-coding-przewodnik). Podłoga rośnie dla wszystkich. **Agentic engineering trzyma SUFIT.** To realna dyscyplina inżynierska: utrzymać profesjonalną poprzeczkę jakości - żadnych nowych podatności, *to wciąż twój* software, ty za niego odpowiadasz - **jednocześnie** przyspieszając. Agenci to „spiky entities": zawodni, trochę stochastyczni, ale ekstremalnie potężni. Sztuką jest ich koordynować, nie obniżając przy tym poprzeczki. > „Vibe coding is about raising the floor for everyone... agentic engineering is about preserving the quality bar of what existed before in professional software." *(Vibe coding podnosi podłogę dla każdego; agentic engineering chroni poprzeczkę jakości, która istniała w profesjonalnym software'rze.)* I tu jest druga strona medalu. Sufit jest bardzo wysoko. Stary „10× engineer" jest teraz wzmocniony *daleko* poza 10× - dla tych, którzy są w tym dobrzy. To nie jest gra o zerowej sumie między „AI" a „człowiekiem". To dźwignia, która rozjeżdża różnicę między kimś, kto umie kierować agentami, a kimś, kto tylko klika „akceptuj". ## Weryfikowalność - nowe ograniczenie Skąd biorą się mocne i słabe strony tych modeli? Karpathy daje na to jedno z najlepszych wyjaśnień, jakie słyszałem, i sprowadza się ono do jednego słowa: **weryfikowalność**. > „The previous generation of computers automated what you could specify. This generation automates what you can verify." *(Poprzednia generacja komputerów automatyzowała to, co umiałeś zaprogramować. Ta generacja automatyzuje to, co umiesz zweryfikować.)* Modele frontierowe trenuje się w gigantycznych środowiskach RL (reinforcement learning, uczenie ze wzmocnieniem) z nagrodami za weryfikację. Dlatego ich możliwości są **jagged** - postrzępione: szczyty tam, gdzie wynik da się zweryfikować (kod, matematyka), doliny gdzie indziej. Jaggedness = to, co *weryfikowalne*, **plus** to, na czym laby akurat zdecydowały się trenować. > „How is it possible that a state-of-the-art model will refactor a 100,000-line codebase or find zero-day vulnerabilities, and yet tell me to walk to a car wash 50 metres away?" *(Jak to możliwe, że topowy model zrefaktoruje kod na 100 tysięcy linii albo znajdzie podatność zero-day, a jednocześnie każe mi iść do myjni oddalonej o 50 metrów?)* Dla foundera płynie z tego konkretny wniosek. Problem **weryfikowalny** jest dziś wykonalny - możesz rzucić w niego RL, zbudować własne środowiska RL albo zrobić fine-tune. To jest twoja fosa, niezależna od tego, co akurat priorytetyzują wielkie laby. Jeśli umiesz tanio i automatycznie sprawdzić, czy wynik jest dobry, masz dźwignię, której konkurencja bez tej weryfikacji nie ma. ## Czego nie da się outsourcować To jest emocjonalny rdzeń całej tezy. Skoro agent zrefaktoruje 100 tysięcy linii, to co właściwie zostaje po stronie człowieka? Zostaje **smak, osąd i nadzór**. Ty trzymasz spec, plan i top-level design - agent „wypełnia luki". Zostają też **fundamenty ponad trivia API**. Karpathy nie pamięta już, czy to `keepdim` czy `axis`, `reshape` czy `permute` - „stażysta" ma do tego idealną pamięć. Ale ty wciąż musisz rozumieć, co dzieje się pod spodem (np. jak działają widoki tensora i pamięć), żeby nie kopiować po cichu danych w niewłaściwym miejscu. A na samym dnie jest jedna rzecz, której nie oddasz nigdy: > „You can outsource your thinking, but you can't outsource your understanding." *(Myślenie możesz oddelegować. Zrozumienia - nie.)* Karpathy mówi wprost, że czuje się teraz **wąskim gardłem** - to on wie, co zbudować i dlaczego, i to on kieruje agentami. I dobrze ilustruje to anegdota o zawodności: agent w MenuGenie próbował dopasowywać użytkowników, korelując **adresy e-mail** ze Stripe i Google, zamiast użyć trwałego ID użytkownika. To błąd typu „dlaczego miałbyś to w ogóle zrobić" - taki, który musi wyłapać człowiek, bo agent nie ma osądu, żeby zauważyć, że to bez sensu. To samo widzę u siebie: agent przyspiesza wszystko, ale to ja jestem reviewerem, który decyduje, co jest poprawne. Zbudowanie własnej bazy wiedzy, którą agent utrzymuje, to dokładnie inwestycja w **zrozumienie** - opisałem ten proces w tekście o [bazie wiedzy w stylu Karpathy'ego](/blog/llm-knowledge-base-brain-karpathy). ## Buduj dla agentów, nie dla ludzi Jeśli agenci stają się głównym „użytkownikiem" twoich systemów, to logiczna konsekwencja jest taka, że trzeba zacząć projektować *dla nich*. Karpathy ma z tym swój prywatny problem (pet peeve - rzecz, która szczególnie go drażni): dokumentacja wciąż jest pisana dla ludzi. „Dlaczego ludzie wciąż mówią mi, co mam zrobić? Co jest tą rzeczą, którą mam wkleić mojemu agentowi?" Kierunek jest taki: - Opisuj systemy **najpierw agentom**. - Rozkładaj pracę na **sensory** (czytają świat) i **aktuatory** (działają na świecie). - Trzymaj struktury danych **legible** - czytelne dla LLM-a. Docelowo zmierzamy do świata, w którym „mój agent rozmawia z twoim agentem" - żeby np. umówić spotkanie. A skoro tak, to nawet **rekrutacja** wymaga przebudowy. Zamiast łamigłówek algorytmicznych dajesz kandydatowi duży projekt - np. „zbuduj bezpiecznego Twitter-clone'a dla agentów, a potem niech 10 agentów próbuje go złamać" - i obserwujesz, *jak* włada narzędziami. Bo to władanie narzędziami, a nie recytowanie z pamięci, jest teraz prawdziwą kompetencją. ## Nie zwierzęta, lecz duchy Na koniec metafora, która porządkuje całe nastawienie. Karpathy mówi, że nie budujemy zwierząt - *przywołujemy duchy*. > „We're not building animals. We are summoning ghosts." *(Nie budujemy zwierząt. Przywołujemy duchy.)* To statystyczne obwody symulacji - substrat z pre-treningu z doklejonym RL - a nie inteligencje ukształtowane przez ciekawość i ewolucję. Krzyczenie na nie nic nie da. Wartość tej metafory jest praktyczna: to *nastawienie*. Bądź odpowiednio podejrzliwy i sprawdzaj empirycznie, co działa, zamiast zakładać, że model „rozumie" tak jak człowiek. I tu wracam do ciebie. **W którym paradygmacie jest twój produkt** - 1.0, 2.0 czy 3.0? A kiedy wielkie LLM-y będą umiały zrobić twoje zadanie wprost, **co zostanie twoją fosą?** Karpathy zostawia cztery odpowiedzi: 1. **Dane** - zastrzeżone zbiory, własne środowiska RL, model wytrenowany na *twojej* wiedzy i stylu. 2. **Prompty i kontekst** - wypracowane systemy kontekstu i baza wiedzy, którymi sterujesz silnikiem. 3. **System design** - sensory, aktuatory, UX i pętle weryfikacji wokół modelu. Silnik to tylko silnik; ty projektujesz samochód. 4. **Zaufanie** - marka i odpowiedzialność za wynik. „Wytrenowałem LLM na całej wiedzy Elona" brzmi inaczej od przypadkowej osoby, a inaczej od samego Elona. Tego model nie kupi. Jeśli twoja przewaga nie siedzi w żadnej z tych czterech fos, to warto zadać sobie to niewygodne pytanie wcześniej niż później.

Chcesz przebudować swój produkt na agentic engineering?

Pomogę Ci ocenić, w którym paradygmacie działasz, gdzie jest Twoja fosa i jak wdrożyć agentów bez obniżania poprzeczki jakości - od architektury kontekstu po pętle weryfikacji.

Umów bezpłatną konsultację
## Przydatne zasoby - [Andrej Karpathy: From Vibe Coding to Agentic Engineering](https://www.youtube.com/watch?v=96jN2OCOfLs) - Sequoia Capital, 29:49. Źródło pierwotne całej tezy: vibe→agentic, weryfikowalność, jagged intelligence. - [Kaparthy revealed the most profitable business to build (Software 3.0)](https://www.youtube.com/watch?v=hJNp9RwK-Uw) - Dream Labs AI, 14 min. Ujęcie biznesowe paradygmatów i czterech fos. - [Vibe coding: przewodnik](/blog/vibe-coding-przewodnik) - jak podnieść podłogę, czyli druga strona tego rozróżnienia. - [Baza wiedzy w stylu Karpathy'ego](/blog/llm-knowledge-base-brain-karpathy) - jak budować fosę danych i kontekstu. ## FAQ
### Czym różni się Software 3.0 od Software 1.0 i 2.0 według Karpathy'ego? Software 1.0 to pisanie jawnego kodu reguła po regule, Software 2.0 to programowanie przez kurowanie danych i trenowanie wag sieci neuronowej, a Software 3.0 to **prompting** - sterowanie LLM-em przez okno kontekstu. W 3.0 to, co wpisujesz w context window, staje się twoją dźwignią nad modelem traktowanym jak interpreter. Praktyczna konsekwencja: zamiast budować narzędzie, projektujesz system wokół silnika-LLM, który dostarcza gotowy wynik.
### Czym różni się vibe coding od agentic engineering? Vibe coding „podnosi podłogę" - sprawia, że każdy może zbudować software, co jest demokratyzujące. Agentic engineering „trzyma sufit" - to dyscyplina inżynierska polegająca na utrzymaniu profesjonalnej poprzeczki jakości (brak nowych podatności, odpowiedzialność za kod) przy jednoczesnym przyspieszeniu z pomocą agentów. Krótko: vibe coding to dostępność, agentic engineering to jakość pod presją tempa.
### Co to jest jagged intelligence w modelach AI? Jagged intelligence (postrzępiona inteligencja) to zjawisko, w którym model ma szczyty możliwości w domenach weryfikowalnych, jak kod czy matematyka, i doliny w innych zadaniach. Wynika to z treningu w środowiskach RL z nagrodami za weryfikację - model jest świetny tam, gdzie wynik da się automatycznie sprawdzić. Stąd paradoks: ten sam model zrefaktoruje 100 tysięcy linii kodu, a pomyli się przy prostym zadaniu z życia codziennego.
### Czego nie da się outsourcować agentom AI przy budowie produktu? Nie da się oddelegować zrozumienia, osądu i nadzoru - jak ujmuje to Karpathy, „myślenie możesz oddelegować, zrozumienia nie". Człowiek zostaje właścicielem specyfikacji, planu i top-level designu, a agent wypełnia luki. Liczą się też fundamenty (rozumienie, co dzieje się pod spodem), bo agent potrafi popełnić bezsensowny błąd, którego sam nie wychwyci.
### Jak budować produkt „agent-native", czyli dla agentów, a nie tylko dla ludzi? Opisuj systemy najpierw agentom (np. blok instrukcji do wklejenia, a nie dokumentacja dla człowieka), rozkładaj pracę na sensory czytające świat i aktuatory działające na świecie oraz trzymaj dane w formie legible dla LLM. Celem jest świat, w którym „mój agent rozmawia z twoim agentem". Warto też przebudować rekrutację: zamiast łamigłówek dawaj realny projekt i obserwuj, jak kandydat włada narzędziami.
--- # Pre-revenue startup bez działu finansów. Anatomia systemu agentów AI. Source: https://pawel.lipowczan.pl/blog/system-agentow-ai-skills-rules-kontekst Published: 2026-05-09 Dziś jest 9 maja 2026. Dwa dni temu prowadziłem prelekcję na NoCode Poland #4 - *"Jak 2 założycieli robi robotę całego zespołu"*. Tytuł brzmi jak clickbait, ale tydzień wcześniej, 1 maja, część tego, o czym mówiłem, **jeszcze nie istniała**. Pierwszego maja w 200IQ LABS - pre-revenue spółce, którą prowadzę z Przemkiem - nie było działu finansów. Zero raportu zarządczego za kwiecień. Zero kontroli kosztów na poziomie kategorii. Zero spec-a. **Trzeciego maja był pierwszy close kwietnia**: 30+ reguł klasyfikacji wytrenowanych z zera, raport zarządczy w `monthly/2026-04.md`, real EBITDA −16 804 PLN udokumentowane z trzema decyzjami korekcyjnymi planu na kolejne miesiące. Dwa dni od zera do działającego systemu zarządczego. Powiem szczerze - robiłem to częściowo na potrzeby prelekcji. **Event Driven Development™** (eventem była ta prelekcja). Ale ten system od dawna miał powstać. Wolę wdrażać PR-y na produkcję niż siedzieć nad numerami - taka prawda founderska. Prelekcja po prostu wymusiła timing. Ten artykuł nie jest o tym, że *"AI zastąpi księgowych"*. Nie zastąpi - formalna księgowość przez inFakt + księgową robi compliance, i nadal to robi. Ten artykuł jest o tym, **jak architektonicznie zbudować system agentów AI, który realnie zastępuje role specjalistów w warstwie zarządczej** - w tym, gdzie dział nie istnieje, bo firma jest za mała, żeby zatrudnić, i za duża, żeby improwizować. Trzy filary, do których wracam w całym tekście: 1. **Skills automatyzują procesy.** Workflow staje się rzeczownikiem - `/finances close 2026-04` zamiast 6 manualnych kroków. 2. **Determinizm przez reguły.** Hybryd rules-first + LLM fallback + learning loop. System konwerguje od probabilistycznego do deterministycznego. 3. **Wspólny kontekst.** CLAUDE.md, MEMORY.md, struktura `context/`. Bez tego subagenty są izolowanymi czatami. Z tym - system, który pamięta firmę. ## Architektura: jeden diagram, trzy warstwy ![Architektura systemu agentów AI w 200IQ LABS - główny agent (Claude Code orchestrator) → 5 subagentów (CFO, Marketing, Legal, Tax Advisor, Business Consultant) → warstwa narzędzi (Python tools + External MCP) → wspólny kontekst (CLAUDE.md, MEMORY.md, context/)](/images/architecture-system-agentow-ai.webp) Ten diagram to spine artykułu. Każda z trzech warstw odpowiada jednemu z filarów, do których będę wracał. **Główny agent** to Claude Code działający jako orchestrator. Z jego perspektywy ja jestem użytkownikiem, on jest punktem wejścia. Routing do subagentów dzieje się przez skille - wpisuję `/finances close 2026-04`, agent ładuje skill `finances`, persona przełącza się na CFO. **Subagenty wyspecjalizowane** - CFO, Marketing, Legal, Tax Advisor, Business Consultant - to nie osobne procesy ani osobne modele. To **kombinacje skill + persona + scoped context**. CFO czyta `context/finances/`, Marketing czyta `context/qamera/` (bo Qamera jest naszym produktem) i `context/brand/`, Legal czyta `context/legal-entities.md` i `context/company.md`. Każda rola widzi tylko to, co jej potrzebne - higiena kontekstu. **Warstwa narzędzi** to hybryda Python (default dla zaplanowanego workflow) + External MCP (dla ad-hoc lub gdy API wymusza). Stripe, Revolut, Airtable to Python wrappery w repo - bo wiemy z góry, jak z nich korzystamy. inFakt to MCP, bo endpoint do pobierania kosztów zwraca 403, gdy spółka jest obsługiwana przez biuro księgowe - oficjalny MCP server omija to ograniczenie. Qamera AI ma własne MCP, z którego korzystają zarówno cudzy agenci, jak i my sami - bo do ad-hoc eksploracji własnego produktu MCP jest po prostu szybsze niż pisanie skryptu pod każde nowe zapytanie. **Wspólny kontekst** to substrate całego systemu. CLAUDE.md ładuje się na każdą rozmowę - instrukcje projektu, konwencje, ścieżki. MEMORY.md jest persistent między sesjami. `context/` to strukturalna baza wiedzy: finanse, klienci, projekty, operacje. Subagenty czytają z tej bazy - i dzięki temu nie są izolowanymi czatami. Teraz po kolei. ## Pillar 1: Skills automatyzują procesy **Skill to jednostka, w której zamykasz proces** - trigger, kroki, decyzje, integracje. Przed skillem masz workflow zdefiniowany w głowie i w czyimś notatniku. Z skillem masz workflow jako rzeczownik: `/finances close YYYY-MM`, `/ingest`, `/slides:new`. Komenda zamiast pamięci. Brzmi banalnie. Nie jest. Bez skilla każdorazowo improwizuję: *"hej, zaciągnij dane Stripe za kwiecień, potem Revolut, sklasyfikuj transakcje, sprawdź accruals, zapisz raport"*. To znaczy 6 razy dziennie wymyślam tę samą sekwencję. Z `/finances close 2026-04` - jedna komenda, deterministyczna sekwencja faz. ### `/finances close` - 6 faz ![Sześć faz miesięcznego close w skill `/finances`: PULL → CLASSIFY → REVIEW → ACCRUALS CHECK → COMMIT → REGENERATE. PHASE 3 i 5 to human-in-the-loop](/images/diagram-close-process.webp) ```text PHASE 1: PULL (~2 min) Revolut + Stripe (Python) + inFakt (MCP) + tech-stack PHASE 2: CLASSIFY (~30s) rules.yaml deterministic → LLM fallback PHASE 3: REVIEW (~10-15min interactive) [a/c/r/s] learning loop PHASE 4: ACCRUALS CHECK (~2 min) accrued-liabilities.yaml matching PHASE 5: COMMIT (~10-20 min) mandatory narrative blokuje close PHASE 6: REGENERATE (~30s) 3 dashboards, AUTO:START/END markers ``` Każda faza jest **idempotentna** - restart po awarii nie duplikuje danych. Możesz ten close uruchomić, przerwać po PHASE 3, wrócić za godzinę i kontynuować od miejsca przerwania. To nie jest pipe-and-pray; to operacja transakcyjna w sensie "można powtórzyć i nic złego się nie stanie". ### Wideo: jeden close na żywo

Pełny close kwietnia 2026 uruchomiony jako /finances close 2026-04. ~7 minut, 6 faz, 2 iteracje QA zostały w nagraniu - meta-przekaz: można korygować rozmową z agentem.

W nagraniu zostały **dwie iteracje QA** - momenty, w których zatrzymałem agenta, zauważyłem coś, co nie pasowało, i poprosiłem o korektę bez restartowania całej fazy. Świadomie ich nie wyciąłem. To pokazuje rzecz, która ginie w marketingowych demo: workflow z agentem nie jest skryptem jednoprzejściowym. Można korygować rozmową, agent dopasowuje stan, kontynuuje. ### Generalizacja: każdy proces może być skillem `/finances close` to flagowy skill, ale ta sama mechanika działa wszędzie. `/ingest` klasyfikuje pliki wrzucone do `inbox/` - opisałem to dokładniej w [PIT-38 case study](/blog/pit-38-claude-code-case-study). `/slides:new` generuje szkielet prezentacji. Meta-twist: te slajdy NCP4, o których piszę w intro, powstały tym samym workflow. **Sam skill `/blog-article-writer` używany do tego artykułu** ma własne fazy: prime → plan → execute → validate → translate. ### Mandatory narrative blokuje close PHASE 5 wymaga, żeby człowiek napisał trzy sekcje narracji w `monthly/.md`: 1. Co się działo niespodziewanego (vs plan). 2. Decyzje podjęte w trakcie miesiąca. 3. Plan na następny miesiąc / korekta planu. **Pusty narrative blokuje close.** System nie zamknie miesiąca dopóki nie wpiszesz czegoś w każdą sekcję. To wygląda jak zbędny formalizm - i z perspektywy *"chcę szybko zamknąć"* tak jest. Ale z perspektywy *"za pół roku patrzę na variance EBITDA −3160 vs plan i nie wiem dlaczego"* - to jedyny sposób, żeby decyzje miały trwałą pamięć. **Refleksja jest częścią procesu, nie opcją.** To pattern szerszy niż finanse: za każdym razem, gdy automatyzujesz coś, w czym wartość leży w **rozumieniu** a nie w **wykonaniu** - dodaj human-in-the-loop checkpoint, który nie da się ominąć. Inaczej zautomatyzujesz wykonanie i zniszczysz rozumienie. ## Pillar 2: Determinizm przez reguły Klasyfikacja transakcji jest klasycznym problemem cold-start. Pure rules: każdy nowy vendor wymaga manual setup, miesiąc bez nowych vendorów to dobry tydzień, miesiąc z 5 nowymi to wieczór. Pure LLM: niedeterministyczny, niedrogi w skali, ale niedopuszczalny w finansach - audyt, regulator, przyszła kontrola. Trzecia droga to hybryd: **rules-first + LLM fallback + learning loop**. ### Rules-first ```yaml # rules.yaml - fragment - id: gworkspace-via-gcp pattern: memo: "GCPLD" source: infakt amount_range: [200, 500] classify: category: opex/saas unit: null note: "Google Workspace billowane przez Google Cloud Poland (300-330 PLN/mc)" created_at: 2026-05-04 - id: gcp-infakt pattern: memo: "GCPLD" source: infakt amount_range: [500, 999999] classify: category: cogs/ai-generation unit: qamera note: "Faktury GCP w inFakt z prefiksem GCPLD i kwotą >500 PLN = compute" created_at: 2026-05-04 ``` Najmocniejszy detal: **dwie reguły dla tego samego vendora** (`GCPLD` w inFakt), różniące się tylko `amount_range`. 200-500 PLN → Google Workspace (kategoria `opex/saas`). >500 PLN → compute GCP dla pipeline'u Qamery (kategoria `cogs/ai-generation`). To jest typ niuansu, który LLM zgaduje **zawsze inaczej**. Jednego dnia zaklasyfikuje cały `GCPLD` jako Workspace, drugiego jako compute, trzeciego rozdzieli losowo. Tu mamy odpowiedź zaszytą jako kod. Audytowalne, deterministyczne, idempotentne. Pole `note:` jest równie ważne jak sam pattern. Bez tej notki za pół roku patrząc na regułę nie wiem, dlaczego rozdzielenie po `amount_range` ma sens. **Reguła bez note-a to długofalowy debt.** ### LLM fallback z few-shot examples Gdy żadna reguła nie matchuje, agent leci do LLM z few-shot promptem. Każdy example w `examples.yaml` ma konkretny `reasoning`: ```yaml # examples.yaml - fragment - transaction: date: 2026-04-10 amount: -127.40 memo: "BYTEPLUS API USAGE PAY-AS-YOU-GO" source: revolut classification: category: cogs/ai-generation unit: qamera reasoning: "Byteplus to provider Kling AI używany do generacji video w produkcie" ``` LLM dostaje 15-25 takich przykładów per kategoria. Zwraca klasyfikację plus własne `reasoning`. **`reasoning` widoczny dla użytkownika** w fazie REVIEW - to nie black box. Widzę nie tylko "kategoria X", ale dlaczego. ### Learning loop: REVIEW [a]ccept / [c]hange / [r]ule / [s]kip Pętla per LLM-classified transakcja. Cztery decyzje: | Opcja | Co robi | |---|---| | `[a]ccept` | LLM zgadł, ale to jednorazowy przypadek | | `[c]hange` | LLM się pomylił, ręczna korekta | | `[r]ule` | Akceptuje + tworzy regułę dla podobnych | | `[s]kip` | Nie wiem, wracamy w następnym close | **Złota zasada:** jeśli transakcja prawdopodobnie się powtórzy → `[r]ule`. 30 sekund teraz, oszczędzasz minuty przez kolejne miesiące. Każde `[r]ule` zamraża LLM-owe zgadywanie w deterministycznej regule w `rules.yaml`. Następny close - ta sama transakcja wpada w "auto-classified" deterministycznie. **Real number:** pierwszy close (kwiecień 2026) wytworzył **30+ reguł z zera**. Trzy typy patternów ujawniły się natychmiast: 1. **Po memo** - czyste pattern matching. Byteplus, Cursor, Meta Pay. 2. **Po prefiksie faktury inFakt + kwocie** - `GCPLD` z `amount_range: [200, 500]` to Workspace, z `[500, ∞]` to compute. Ten sam vendor, dwie kategorie. 3. **Po reseller pattern** - Paddle 24 EUR ≈ 100 PLN to n8n cloud. Ale Paddle z inną kwotą = REVIEW (może być inny produkt sprzedawany przez Paddle). System konwerguje od probabilistycznego do deterministycznego. To rzadki design - większość narzędzi do kategoryzacji budżetu jest albo w 100% deterministyczna (sztywne reguły, dużo manuala), albo w 100% probabilistyczna (LLM/ML, brak audytu). Hybryd z `[r]ule` jako pomostem między obiema stronami daje cold-start LLM-a i długoterminowy audit deterministycznych reguł. ### Idempotent regeneration jako ortogonalny mechanizm determinizmu Drugą warstwą determinizmu są dashboardy. PHASE 6 regeneruje 3 pliki: `_dashboard.md`, `_alerts.md`, `_runway.md`. Pytanie: co z ręcznymi notatkami, które dopisałem w trakcie miesiąca? Jeśli regeneracja nadpisuje cały plik - moje notatki znikają i przestaję pisać notatki. Rozwiązanie: markery zasięgu. ```markdown # 200IQ LABS - Finances Dashboard ## Cash position & runway | Metryka | Wartość | |---|---| | Cumulative shareholder loans | 74 300 PLN | | Confirmed financing (czerwiec) | +100 000 PLN | | Accrued liabilities (UoD payout maj) | −24 000 PLN | | Run-rate burn | ~14-14.5k PLN/mc | ## Notatki ręczne (poza markerami - przetrwają regenerację) - 2026-05-04: pierwszy close systemu. Manual regeneracja dashboardów. - Plan finalize: nie wykonano - tryb `draft` świadomie do kolejnego close. ``` Markery `AUTO:START` / `AUTO:END` ograniczają zasięg regeneracji. Wszystko poza markerami **przeżywa każdą regenerację**. Bez tego nikt nie pisze ręcznych notatek w auto-generowanych dokumentach - i te dokumenty stają się martwe. Ten prosty pattern decyduje o tym, czy dashboardy są używane, czy ignorowane. ## Pillar 3: Wspólny kontekst (substrate) Trzeci filar to warstwa, która sprawia, że wszystkie subagenty są spójne. Bez wspólnego kontekstu CFO i Marketing nie wiedzą o sobie nawzajem - masz 5 izolowanych czatów, każdy ze swoją wersją prawdy. Z wspólnym kontekstem masz **system, który pamięta firmę**. ### Trzy warstwy wspólnego kontekstu - **CLAUDE.md** (per-projekt) - instrukcje, konwencje, ścieżki, "co agent ma robić, jak nie wie". Ładuje się na każdą rozmowę. - **MEMORY.md** (cross-conversation) - auto-memory, persistent między sesjami. Tu siedzi *"użytkownik woli krótkie podsumowania"*, *"projekt X używa OpenSpec"*, *"data dziś to YYYY-MM-DD"*. - **`context/`** (knowledge base) - strukturalna baza per dziedzina. `context/finances/`, `context/qamera/`, `context/clients/`, `context/operations/`. Struktura katalogów (uproszczona): ```text agentic-ai-system/ ├── CLAUDE.md # instrukcje projektu ├── memory/ │ └── MEMORY.md # auto-memory, persistent ├── context/ │ ├── finances/ # ← czyta skill CFO │ ├── qamera/ # ← czyta skill Marketing (revenue + brand) │ ├── clients/ # ← czyta skill Business Consultant │ └── operations/ ├── skills/ │ ├── finances/SKILL.md │ ├── ingest/SKILL.md │ └── slides/SKILL.md └── tools/ ├── stripe/ # Python wrappers ├── revolut/ # Python wrappers └── airtable/ # Python wrappers ``` Każdy subagent ma scope - nie czyta wszystkiego. CFO sięga do `context/finances/` i `context/operations/`, dla revenue ma read-only dostęp do `context/qamera/`. Nie czyta `context/clients/` (to zasięg Business Consultanta), bo nie potrzebuje. **Higiena kontekstu** to drugi multiplikator wartości - bez niej każde pytanie do CFO zaciąga połowę firmy do tokenów. ### Python tools (default) vs External MCP (when forced) > **Reguła: Python skrypt dla zaplanowanego workflow. MCP dla ad-hoc lub gdy API wymusza.** | Integracja | Podejście | Dlaczego | |---|---|---| | Stripe (subskrypcje Qamery) | Python | Wiemy z góry, czego potrzebujemy - wrapper to część specyfikacji | | Revolut (transakcje firmowe) | Python | Stała sekwencja w `/finances close`, własny wrapper | | Airtable (CRM) | Python | Powtarzalne operacje na CRM, własny wrapper | | **inFakt** (księgowość) | **MCP** | Endpoint kosztów zwraca 403, gdy spółkę obsługuje biuro księgowe - MCP **wymuszone** | | **Qamera AI** (nasz produkt) | **MCP** | I my, i cudzy agenci. REST API też dostępne, ale do ad-hoc eksploracji MCP jest po prostu szybszy do podpięcia | **Dlaczego Python > MCP, gdy workflow jest planowany:** 1. **Oszczędność tokenów.** MCP ładuje opisy narzędzi do kontekstu **każdej** rozmowy. Skrypt = zero overhead. Skrypt jest wywołany tylko wtedy, gdy agent zdecyduje, że go potrzebuje. 2. **Determinizm.** Skrypt robi dokładnie to, co napisano. Bez "model zinterpretował narzędzie inaczej w innej rozmowie". 3. **Trywialne tworzenie.** *"Agent, potrzebuję skrypt, który wyciągnie z Stripe wszystkie subskrypcje aktywne na dzień X"* → agent generuje, ja code-review, commit. 15 minut. 4. **Pełna kontrola.** Kod w repo, w wersji, czytelny dla innego dewelopera. Mogę debugować Pythonem, dodać logging, zmienić output format - bez negocjacji z protokołem MCP. **Dlaczego MCP, gdy używamy ad-hoc:** Jeśli dopiero **odkrywam**, jak będę z czegoś korzystać - pisanie skryptu jest przedwczesne. Nie wiem jeszcze, jakie pola mnie interesują, jakie filtry, jaki output format. MCP daje agentowi natychmiastowy dostęp do API; po kilku iteracjach widzę pattern, który warto utrwalić - i wtedy piszę skrypt. **Skrypt jest częścią specyfikacji workflow; MCP jest narzędziem do eksploracji.** Z Qamerą tak właśnie korzystamy: cudzy agenci uderzają w MCP, my też - gdy chcemy szybko sprawdzić coś we własnym produkcie. Gdybyśmy mieli stały workflow w stylu *"co tydzień generuj X dla klienta Y"* - pisalibyśmy skrypt. Na razie nie mamy, więc MCP wystarcza. Z inFaktem nie mamy wyboru - REST API zwraca 403 na endpoint kosztów, gdy spółkę obsługuje biuro księgowe. MCP tu nie jest wyborem architektonicznym, tylko jedyną dostępną drogą do danych. ### Subagenty czytają tylko swój scope Wracając do pierwszego diagramu: każdy subagent (skill + persona) ma scope w swojej linii kontekstu. Skill `finances` ma front-matter, który jawnie deklaruje: ```yaml context_scope: - context/finances/ - context/operations/ - context/qamera/ # read-only, dla revenue z subskrypcji ``` Bez tego ograniczenia agent pyta *"jakie były transakcje w marcu"* i ładuje kontekst całej firmy do tokenów. Z ograniczeniem - ładuje tylko `context/finances/transactions/2026-03.yaml` plus reguły. Mniej tokenów, szybsza odpowiedź, mniejsze ryzyko halucynacji. Pisałem o tym samym wzorcu w innym kontekście - separacja `inbox/` / `archive/` / `data/` w [PIT-38 case study](/blog/pit-38-claude-code-case-study). Wzorzec się generalizuje: **agent pracuje na czystych, przetworzonych plikach, nie na surowiznie**. ## Case study: real numbers z kwietnia 2026 Dowód, że to nie demo. Pierwszy close, wykonany 2026-05-04, wygenerował: | Linia | Actual | Plan | Δ | |---|---:|---:|---:| | Revenue | 347 PLN | 598 PLN | −42% | | Costs | 17 151 PLN | 14 242 PLN | +20% | | EBITDA | −16 804 PLN | −13 644 PLN | −23% | | `cogs/ai-generation` | 1 709 | 700 | +144% 🔴 | | `opex/marketing` | 2 996 | 1 500 | +100% 🔴 | | `opex/saas` | 2 112 | 1 633 | +29% 🟡 | **Trzy decyzje wynikające z close-a** (zapisane w `monthly/2026-04.md`): 1. **GCP +144% vs plan** → plan był błędny (700 PLN), realny run-rate 1000-1600 PLN/mc historycznie. Plan May-Dec podniesiony retroaktywnie do 1500-1800 PLN. Świadoma decyzja **NIE optymalizować GCP** - część tego to content marketingowy (own + materiały dla klientów). Klasyfikacja "to faktycznie marketing" nie zmienia kategorii (GCP zostaje w `cogs/ai-generation`), ale uzasadnia wyższy plan. 2. **Cursor billowany 4×/mc** (716 PLN vs plan 217) → planowana migracja **Cursor → Claude** (sztywny ~90 EUR/mc cap, predictable koszt zamiast usage-based). 3. **Meta Ads off od maja** - strategia była błędna, kierunek do dopracowania. Plan kontroli: w close maja sprawdzimy, że spend < 300 PLN. **Pointa:** decyzje wynikają **z systemu, ale nie podejmuje ich system**. System pokazuje variance i wymaga w PHASE 5 narrative. Decyzję podejmuję ja - ale mam ją udokumentowaną na wypadek, gdy za pół roku zapomnę dlaczego. Plan break-even: wrzesień 2026 (revenue 12 184 vs costs 12 842, EBITDA −658). Plan agresywny - wymaga walidacji w kolejnych close-ach. **Plan = żywy dokument**, korygowany po każdym close gdy run-rate się zmienia. Drift > 30% = wymaga rewizji. ## Czego nie polecam Mapy ryzyk, mirror sekcji z PIT-38 case study. - **Nie polecam tego setupu osobie bez programistycznego komfortu.** Markdown, YAML, git, terminal, edycja convention files - wymaga rozumienia narzędzi. Bez tego strata czasu na setup zje wszystkie zyski. Bezpieczniejsza ścieżka: gotowy SaaS do controllingu (Causal, Runway, Pry) + Excel + księgowa. - **Manual first, automate after pain.** Pierwsze 2-3 close-y MUSZĄ być manualne. Bez ręcznego wykonania połowa reguł byłaby zaprojektowana źle. Pierwszy close ujawnił 30+ patternów do zakodowania - w tym nieoczywiste (np. że `GCPLD` to dwie kategorie, nie jedna). **Skrypty piszesz po, nie przed.** - **System zarządczy ≠ księgowość formalna.** inFakt + księgowa robią compliance. Ten system robi decyzje. Nie podmieniaj jednego za drugie - uzupełniają się, nie zastępują. Ja nadal płacę księgowej comiesięczny ryczałt; ten system nie usuwa tej linii z budżetu. - **LLM-y robią błędy arytmetyczne.** Sumowanie do silnika obliczeniowego (Excel, kalkulator, system MF). LLM ma wartość w **strukturze i interpretacji**, nie w arytmetyce. Pisałem o tym w [PIT-38 case study](/blog/pit-38-claude-code-case-study) - zostawiłem 6 groszy na stole, bo LLM pomylił się przy odejmowaniu dwóch sześciocyfrowych liczb. - **Nie skaluje się liniowo poza 1 firmę.** System zaprojektowany dla 200IQ LABS. Wzorce się generalizują (skills, rules, kontekst); konkretne YAML-e nie. *"Nie kopiujcie setupu, wyciągajcie zasady"* - to było motto talku NCP4 i jest motto tego artykułu. ## Kluczowe wnioski 1. **Skills sprawiają, że proces staje się rzeczownikiem.** `/finances close YYYY-MM` zamiast 6 manualnych kroków. Każdy powtarzalny workflow zasługuje na własny skill - z fazami, z idempotency, z mandatory checkpoints tam, gdzie wartość leży w rozumieniu. 2. **Determinizm budujesz iteracyjnie.** Rules-first + LLM fallback + opcja `[r]ule` w REVIEW = system konwerguje z probabilistycznego do deterministycznego. Po pierwszym close mieliśmy 30+ reguł. Po dziesiątym będziemy mieli ~150, LLM wywoływany sporadycznie. 3. **Wspólny kontekst (CLAUDE.md, MEMORY.md, `context/`) to substrate, nie ozdoba.** Bez niego subagenty są izolowanymi czatami. Z nim - system, który pamięta firmę i dzieli wiedzę między rolami. 4. **Python skrypt dla zaplanowanego workflow, MCP dla ad-hoc.** Skrypt jest częścią specyfikacji - token-efficient, deterministic, code-reviewable. MCP daje natychmiastowy dostęp gdy nie wiesz jeszcze, jak będziesz korzystać, lub gdy API wymusza. 5. **Refleksja jest częścią procesu, nie opcją.** Mandatory narrative blokuje close. Bez tego za pół roku nie wiesz dlaczego coś było. Wolę wdrażać kod niż liczyć - i agenty pomagają mi w obu. Ten system od dawna miał powstać. Eventem napędzającym była prelekcja. Działa.

Prowadzisz pre-revenue startup i myślisz o systemie zarządczym AI?

Pomagam founderom i konsultantom technologicznym projektować architekturę agentów AI dla operacji firmy - skills, deterministyczne reguły, wspólny kontekst. Pokażę, jak taki setup mógłby wyglądać u Ciebie.

Umów bezpłatną konsultację
## Przydatne zasoby - [Claude Code - dokumentacja](https://docs.claude.com/en/docs/claude-code/overview) - referencja do skills, settings, MCP, hooks - [Model Context Protocol - spec](https://modelcontextprotocol.io/) - czym jest MCP i kiedy ma sens - [PIT-38 case study](/blog/pit-38-claude-code-case-study) - ten sam wzorzec architektoniczny w kontekście rozliczeń podatkowych - [Skills 2.0 - multi-agent system do zarządzania firmą](/blog/skills-2-0-multi-agent-system-zarzadzanie-firma) - szerszy kontekst Skills 2.0 i rola subagentów - [OpenSpec workflow - strukturyzowana praca z AI](/blog/opsx-workflow-strukturyzowana-praca-z-ai) - spec-driven development jako fundament pod skille - [Spec-driven SEO na portfolio i Qamera AI](/blog/spec-driven-seo-portfolio-qamera-ai) - inny case study, ten sam typ workflow ## FAQ
### Czy ten wzorzec (skills + rules + wspólny kontekst) działa tylko w Claude Code, czy w innych agentach też? Wzorzec jest agent-agnostic. Skills mapują się na komendy/funkcje w innych systemach (Cursor commands, n8n workflows, custom CLI). Rules + few-shot examples to standardowa technika ML. Wspólny kontekst (markdown + struktura katalogów) działa wszędzie, gdzie agent ma dostęp do plików. **Konkretne mechanizmy** (`CLAUDE.md`, MCP, `MEMORY.md`) są specyficzne dla Anthropic, ale **zasady architektoniczne** się generalizują na każdy LLM-driven system.
### Ile czasu zajmuje zbudowanie takiego systemu od zera? Pierwszy działający close - z OpenSpec, schemami YAML i 30+ regułami - zajął nam dwa dni intensywnej pracy w dwie osoby. Pełna stabilizacja (auto-pull skrypty, scheduled close, integracja z formalną księgowością) zakłada się na 4-6 tygodni przy regularnej miesięcznej kadencji. **Każdy dzień, w którym używasz manualnie zaprojektowanego skilla, jest dniem produktywnym** - nie czekasz, aż system będzie kompletny, bo wartość pojawia się natychmiast po pierwszym close-u.
### Dlaczego Python skrypt zamiast MCP, jeśli MCP jest standardem? MCP ładuje opisy narzędzi do kontekstu **każdej** rozmowy - to koszt tokenów i potencjalna niejednoznaczność interpretacji w różnych sesjach. Python skrypt to zero overhead, deterministyczny output, kod w repo do code-review. **Skrypt piszemy dla zaplanowanego workflow** (Stripe, Revolut, Airtable - wiemy z góry, czego potrzebujemy, więc skrypt jest częścią specyfikacji). **MCP używamy dla ad-hoc** (eksploracja własnej Qamery, gdy nie wiemy jeszcze jakie zapytania będą się powtarzać) **lub gdy API wymusza** (inFakt zwraca 403 na endpoint kosztów, gdy spółkę obsługuje biuro księgowe). Pattern: po kilku iteracjach ad-hoc widać, co warto utrwalić w skrypcie.
### Czy LLM nie pomyli się przy klasyfikacji finansowej? Co z audytem dla regulatora? Hybryd (rules-first + LLM fallback) projektujemy specjalnie po to, żeby zminimalizować LLM-classified transakcje. Każda LLM-classified transakcja przechodzi przez REVIEW - człowiek widzi `reasoning` i decyduje `[a]ccept`/`[c]hange`/`[r]ule`/`[s]kip`. Po pierwszym close mieliśmy 30+ reguł, które deterministycznie wyłapują 80%+ kolejnych transakcji. **Audit trail leży w `rules.yaml` plus `monthly/.md`** - każda decyzja udokumentowana, każda reguła ma `note:` z uzasadnieniem. Formalna księgowość (compliance) idzie nadal przez inFakt + księgową; ten system jest zarządczy, nie compliance.
### Co robi mandatory narrative, jeśli mam awaryjny close i nie zdążę napisać 3 sekcji? Trzy sekcje (niespodzianki / decyzje / plan korekta) blokują close-a do momentu wypełnienia - to świadoma decyzja. W skrajnym przypadku możesz wpisać do każdej sekcji jedno zdanie (*"brak niespodzianek vs plan"*, *"brak decyzji"*, *"kontynuujemy bez korekt"*). System nie ocenia jakości narrative, tylko jego obecność. **Punkt jest taki, że za pół roku patrząc na variance EBITDA potrzebujesz dowolnego kontekstu - lepiej krótkiego niż żadnego.** Ten formalizm jest celowy.
### Czy ten system działa dla firmy na revenue, nie tylko pre-revenue? Tak, wzorzec się skaluje - kategorie P&L i caps trzeba dopasować, mandatory narrative staje się jeszcze ważniejszy (więcej transakcji = więcej decyzji do udokumentowania), learning loop daje większy ROI (więcej powtarzalnych vendorów). **Ale w post-revenue spółce z dużym wolumenem prawdopodobnie potrzebujesz pełnoprawnego ERP** - system zarządczy z markdownem i YAML-em ma sufit przy ~setkach transakcji/miesiąc i kilku osobach decyzyjnych. **Nasz scope:** pre-revenue → early-revenue, 1-5 osób w firmie. Powyżej tego skali architektura wymaga twardszych narzędzi.
--- # Miałem 3 dni do PIT-38 bez księgowej. Wystarczyły 2 godziny. Source: https://pawel.lipowczan.pl/blog/pit-38-claude-code-case-study Published: 2026-04-29 Dziś jest 29 kwietnia 2026. Termin złożenia PIT-38 mija jutro. Przedwczoraj wieczorem wystartowałem od jednego zdania - _"utwórz nowy projekt PIT-38"_. Wczoraj po południu deklaracja była w MF, podatek zapłacony, UPO leży w `output/`. Łącznie **~2 godziny** aktywnej pracy. **5 430 transakcji** z dwóch giełd krypto. **174 895,50 PLN** niewykorzystanych kosztów krypto z 2024 do uwzględnienia. **2 PIT-8C** z brokerów, **1 raport** o dywidendach zagranicznych. Złożone i przyjęte przez MF na **51,5 godziny** przed deadlinem. Końcowa dopłata: **172 PLN**. Zwykle robi to moja księgowa. W tym roku przegapiłem timing - i to się okazało jednym z bardziej pouczających doświadczeń, jakie miałem z agentic workflow. Ten artykuł nie jest o tym, jak nauczyłem się rozliczać PIT-38. Wiedziałem co robić, bo robię to od lat (przez księgową). Jest o tym, jak **dwie godziny aktywnej pracy z dobrze poustawianym agentem** odtworzyły pracę usługi eksperckiej - z lepszym poziomem dokumentacji niż dostawałem od księgowej. ## Zwykle PIT-38 robi mi księgowa PIT-38 składam co roku. Papiery wartościowe, fundusze, krypto, dywidendy zagraniczne. Nigdy sam. Zawsze przez moją księgową - ja podsyłam dane, ona wypełnia, sprawdza, składa. Działa od lat. W tym roku przegapiłem timing. Termin 30 kwietnia, ja zacząłem ogarniać dokumenty 27 kwietnia wieczorem. Księgowa nie weźmie projektu z 3-dniowym buforem - i słusznie. Byłem zmuszony zrobić to sam. Dwa wyjścia: (a) panika i prowizoryczna deklaracja na ostatni moment, albo (b) sprawdzenie, czy ten cały AI-workflow, którym buduję rzeczy dla klientów, może zastąpić księgową w dwa dni. Wybrałem (b). Działało. To zmienia stake'a artykułu. To nie jest "lifehack dla osób, które nie chcą iść do księgowej". To **case study o tym, kiedy automatyzacja AI realnie podchodzi pod usługę ekspercką**, którą do tej pory robił człowiek. ## Architektura projektu: każdy katalog ma jedną odpowiedzialność Wystartowałem od polecenia _"utwórz nowy projekt PIT-38"_. Bez briefu, bez dokumentu wymagań, bez listy zadań. Agent dopytał o rok podatkowy i termin, sięgnął po wewnętrzną konwencję `_template.md` (którą zna z `CLAUDE.md` repo) i postawił strukturę: ```text PIT_38/ inbox/ # drop zone, raw files archive/ # po przetworzeniu (agent NIE czyta) data/ # knowledge - agent reads freely deliverables/ # checklist, blog inputs output/ # finalna deklaracja PDF, UPO project.md # cel, status, decyzje catalog.md # indeks plików ``` Każdy katalog ma jedną odpowiedzialność. Agent nie zagląda do `archive/` ani do `output/`, chyba że jawnie poproszę. To **nie bezpieczeństwo** - to **higiena kontekstu**. Agent pracuje na czystych, przetworzonych plikach `data/*.md`, a nie na 4096 wierszach raw CSV przy każdym pytaniu. Brzmi banalnie. Nie jest. Bez tej separacji każde pytanie typu _"jakie były transakcje w marcu?"_ zaciągnęłoby cały surowy CSV do kontekstu - i albo wybiło mnie z limitu tokenów, albo zmusiło model do "sumowania na palcach". Z separacją: agent dostaje gotowy `data/crypto-summary-2025.md` z 11 pozycjami i pracuje na strukturze, nie na surowiznie. To jest pierwszy multiplikator wartości - i działa **tylko dlatego**, że konwencja jest opisana w `CLAUDE.md`. Bez tego pliku ten projekt zająłby tyle samo czasu co praca ręczna. ![Workflow funnel: 5 niezależnych źródeł danych zbiega się w /ingest, agent klasyfikuje 5430 surowych transakcji do 11 zdarzeń podatkowych, deklaracja trafia do MF z buforem 174 895 PLN przeniesionym na 162 948 PLN](/images/diagram-pit-38-funnel-pl.webp) ## `/ingest`: jedna komenda, cztery kroki Wrzuciłem 7 plików do `inbox/`: dwa PIT-8C (XTB i SFIO), trzy CSV-ki z giełd krypto, raport dywidend, zeszłoroczną deklarację jako referencję historyczną. Komenda: `/ingest PIT_38`. Co dzieje się pod spodem: 1. **Identyfikacja typu pliku** - agent rozpoznaje PIT-8C po strukturze nagłówków, raport krypto po typowych kolumnach (timestamp, asset, type, amount). 2. **Routing** - PIT-8C trafia do sekcji C deklaracji, raporty krypto do sekcji E, dywidendy do sekcji G. Każdy plik dostaje osobny `data/{źródło}-{rodzaj}-2025.md` z normalizacją. 3. **Ekstrakcja** - z PDF-ów wyciągane są kluczowe pozycje (przychód, koszt, podatek u źródła), z CSV-ek klasyfikacja typów transakcji. 4. **Archiwizacja** - surowe pliki idą z `inbox/` do `archive/`, `catalog.md` dostaje update, `project.md` log zmian. Konkretny przykład: PIT-8C od XTB i SFIO oba zostały zsumowane w `pit38-calculation.md` w sekcji C, z numerami pozycji zgodnymi z **PIT-38 wersja (18)** za 2025. Agent sam zauważył, że to nie ten sam wzór co w 2024 - o czym za chwilę. Lekcja na tym etapie: `/ingest` to **nie batch processing**. Wartość jest w pętli iteracyjnej, którą zaraz pokażę. ## Sześć rzeczy, których nie spodziewałem się od agenta To jest serce tego artykułu. Sześć konkretnych rzeczy, które agent zrobił w te dwa dni, a których - gdybym pracował sam z Excelem - albo bym nie zrobił, albo zrobiłbym z błędem. ### 1. Bufor 174 895,50 PLN - auto-pull z PIT-38 2024 W polskim PIT-38 niewykorzystane koszty nabycia krypto przechodzą na kolejne lata bez ograniczenia czasowego (art. 22 ust. 16). Standardowy mechanizm dla aktywnego inwestora w krypto. Ja wiedziałem, że ten bufor istnieje - rozliczam PIT-38 od lat. Ale wiedziałem to z perspektywy _"powiedzieć o tym księgowej"_, nie z perspektywy _"wpisać dokładnie w pozycję 38 nowego formularza"_. Agent przy pierwszym ingest poprosił o zeszłoroczną deklarację. Wrzuciłem PDF z PIT-38 za 2024. Agent sam wyciągnął kwotę z konkretnej pozycji i wpisał w odpowiednie pole nowej deklaracji - tam, gdzie staje się "kosztami z lat ubiegłych". Wartość dla mnie: zazwyczaj ten bufor _"pamięta"_ księgowa, bo ma moją deklarację z poprzedniego roku w swoim CRM. Bez niej musiałbym sam pamiętać o pobraniu PDF, otwarciu go, sprawdzeniu pozycji. **Agent zrobił to za jednym pytaniem.** To jest niedoceniana klasa wartości - automatyzacja "łatwo dostępnej pamięci kontekstu poprzednich lat". Robi to księgowa. Robi to agent z dostępem do historycznych plików w `archive/`. ### 2. PIT-38 wersja (17) vs (18) - drobiazg, który psuje deklarację W PIT-38 za 2024 koszty z lat ubiegłych miały jedną pozycję, w 2025 mają już inną. Sekcja C dostała w wersji (18) dodatkowy wiersz na zwolnienia z art. 21 ust. 1 pkt 105a - i wszystko poniżej przesunęło się o dwa wiersze. Wiele blog-postów i forów internetowych wciąż odnosi się do **starej numeracji**. Gdybym wpisał dane w stare pozycje, deklaracja zostałaby odrzucona albo - gorzej - przyjęta z błędną kalkulacją, która wyszłaby przy kontroli. Agent wyłapał to po wgraniu pustego wzoru PIT-38(18) jako referencji. Porównał strukturę z zeszłoroczną deklaracją, zaznaczył różnice w `data/sources.md`, użył nowej numeracji w finalnej kalkulacji. Lekcja: **zaufaj dokumentom referencyjnym, nie pamięci modelu.** Wgrany aktualny wzór formularza > to, czego model nauczył się z internetu rok temu. ### 3. 5 430 transakcji → 11 zdarzeń podatkowych 4096 transakcji z jednej platformy plus 1334 z drugiej. Razem 5430 raw. Po klasyfikacji: **11 zdarzeń podatkowych**, które trafiają do deklaracji. Reszta to operacje wewnętrzne platform - transfery, korekty techniczne, naliczenia, drobne rozksięgowania, które nie podlegają wykazaniu. Wartość LLM nie jest tutaj w **sumowaniu**. Wartość jest w **klasyfikacji**. Każdy z około 25 typów transakcji w raporcie został przyporządkowany do jednej z trzech kategorii: zdarzenie podatkowe / operacja wewnętrzna / neutralne. Każda kategoryzacja z udokumentowanym uzasadnieniem prawnym w `data/sources.md`. > Nie wchodzę w szczegóły konkretnych typów transakcji ani interpretacji prawnej - to wykracza poza ten case study i wymaga indywidualnej konsultacji z doradcą podatkowym. Pokazuję jedynie skalę pracy klasyfikacyjnej. Bez tej klasyfikacji próba ręcznego rozliczenia 5430 wierszy CSV byłaby albo niemożliwa w dwa dni, albo prowadziłaby do dramatycznego zawyżenia podstawy - gdybym potraktował każdą transakcję jako zbycie. Skala, którą LLM redukuje, jest nietrywialna. ### 4. Kursy NBP D-1 - 10 edge cases z kalendarzem świąt Każdą transakcję w EUR/USD trzeba przeliczyć na PLN po **kursie średnim NBP z dnia roboczego poprzedzającego datę zbycia** (art. 11a ustawy o PIT). Brzmi prosto. W praktyce 10 z 11 transakcji wymagało indywidualnej obsługi, bo wypadły w dni nierobocze: | Data transakcji | Dzień tygodnia | Kurs NBP z dnia | | --- | --- | --- | | 2025-02-15 | sobota | piątek 2025-02-14 | | 2025-03-16 | niedziela | piątek 2025-03-14 | | 2025-05-02 | piątek | środa 2025-04-30 (1 maja = święto) | | 2025-05-18 | niedziela | piątek 2025-05-16 | | 2025-06-15 | niedziela | piątek 2025-06-13 | Agent zrobił to ręcznie dla każdej z 10 transakcji: wyliczył dzień tygodnia, sprawdził święta państwowe (1 maja, 3 maja, Boże Ciało, 15 sierpnia, 1 i 11 listopada, 25-26 grudnia), pobrał kurs z NBP, przeliczył. To jest **dokładnie ten typ pracy, w którym ludzki mózg szybko popełnia błąd** - bo każdy edge case wymaga osobnego sprawdzenia kalendarza. Człowiek robi 1-2 błędy na 10 transakcjach (zapomnienie o 1 maja jest w polskich rozliczeniach częste). Agent zrobił 0. Klasa zadań _"dużo małych mechanicznych decyzji z subtelnymi regułami"_ to mocna strona LLM-a. Nie sumowanie. Nie kreatywność. Systematyczne stosowanie reguł do każdej z 10 sytuacji bez znudzenia i bez pomijania. ### 5. 6 groszy, których LLM nie umie sumować Moja kalkulacja sekcji C: strata **966,98 PLN**. Twój e-PIT po wczytaniu danych: **966,92 PLN**. Różnica 6 groszy. Błąd arytmetyczny po stronie LLM przy odejmowaniu dwóch 6-cyfrowych liczb (16 188,05 − 15 221,13). Lekcja: **LLM-y robią błędy arytmetyczne w 6-cyfrowych dodawaniach.** Zostaw sumowanie kalkulatorowi, Excelowi albo systemowi MF. LLM ma wartość w **strukturze i interpretacji**, nie w arytmetyce. To nie jest porażka - to mapa, gdzie używać tego narzędzia, a gdzie nie. W tym konkretnym przypadku Twój e-PIT poprawił mi błąd za darmo. W innym scenariuszu - gdybym składał papierowo - różnica 6 groszy poszłaby do MF i prawdopodobnie nikt by się nie obraził, ale chodzi o zasadę: **arytmetyka do silnika obliczeniowego, nie do modelu językowego**. ### 6. Ingest jako progresywne odkrywanie - najmocniejszy beat Najbardziej wartościowa rzecz, jaką agent zrobił w te dwa dni, to nie sumowanie 5 tysięcy transakcji. To było **pytanie o rzeczy, których nie miałem na liście**. Nie miałem listy _"co potrzeba do PIT-38"_. Wrzucałem to, co miałem pod ręką. Po każdej iteracji agent mówił: _"OK, to jest, ale brakuje X"_ albo _"uwaga, to oznacza Y, sprawdź Z"_. Kluczowy moment: **dywidendy zagraniczne**. - Nie wiedziałem, że mam dywidendy z 2025 (drobne ETF-y w XTB, łącznie ~958 PLN brutto). - Nie wiedziałem, że trzeba je rozliczać osobno od reszty (sekcja G PIT-38, art. 30a, 19% PL minus podatek u źródła). - Po pierwszej partii ingest agent zapytał: _"a co z dywidendami? PIT-8C ich nie zawiera, XTB wystawia osobny raport"_. - Pobrałem **XTB Raport Dodatkowy do PIT-38** - okazało się, że było 182 PLN podatku 19% PL od dywidend zagranicznych do uwzględnienia. Bez tego pytania złożyłbym deklarację **bez sekcji G**. Skutek: niedopłata 172 PLN, ryzyko kontroli, odsetki. Drugi moment: **historia 2024**. Agent przy pierwszym ingest zapytał, czy mam zeszłoroczną deklarację. Wrzuciłem PDF - i wtedy wyłonił się bufor 174 895,50 PLN, który opisałem wyżej. Bez tego ruchu zapłaciłbym ~2200 PLN podatku zamiast 0. To jest **istota wartości**: nie _"agent zsumował transakcje"_, tylko **"agent wiedział, czego nie wiem"**. Iteracyjne dopytywanie + klasyfikacja każdego dokumentu według taksonomii PIT-38 = niemożliwe do osiągnięcia ręcznie bez specjalistycznej wiedzy podatkowej. Ingest workflow ≠ batch processing. Wartość jest w pętli: wrzucasz → agent klasyfikuje → identyfikuje braki → prosi o dodatkowe dane → wrzucasz znowu. Po 3-4 iteracjach masz komplet, którego sam byś nie zebrał. ## Krótkie wyjaśnienie: jak działa opodatkowanie krypto w PL Drobne wyjaśnienie dla osób spoza tematu krypto-podatków, żeby _"bufor 174k"_ nie zabrzmiał jak _"Paweł stracił 174k"_. W polskim prawie (art. 17 ust. 1 pkt 11 ustawy o PIT) zdarzenie podatkowe powstaje dopiero przy **wymianie krypto na walutę tradycyjną lub na towar**. Dopóki trzymasz pozycję w krypto - ile by się nie zmieniała wartość rynkowa - nic nie wykazujesz. Wymiana krypto-krypto też jest neutralna (art. 17 ust. 1f). Jeśli kupisz krypto za 100 000 zł i nie sprzedasz na fiat przez kilka lat, te 100 000 zł istnieje jako **udokumentowany koszt nabycia** (art. 22 ust. 14) i czeka. W roku, w którym sprzedasz krypto na fiat, ten koszt obniża podstawę opodatkowania. Jeśli koszty > przychody w danym roku - nadwyżka **przechodzi na kolejne lata bez ograniczenia czasowego** (art. 22 ust. 16). Stąd _"bufor 174 895 PLN z 2024 → 162 948 PLN na 2026"_ w moim case'ie. To **nie strata** - to wydatki na zakupy krypto, których jeszcze nie zamknąłem sprzedażą na fiat. Bufor zmniejsza się dopiero wtedy, gdy realnie sprzedaję krypto na PLN/EUR/USD. To jest mocno uproszczony opis mechanizmu. Pełne rozumienie wymaga konsultacji z doradcą - to nie jest porada podatkowa. ## Decyzje interpretacyjne: gdzie człowiek wraca do gry Przy 5400+ transakcjach z różnych platform część kategorii zdarzeń ma **niejednoznaczną kwalifikację podatkową**. Istnieją różne interpretacje KIS, opinie doradców, wyroki sądów administracyjnych dotyczące zbliżonych konstrukcji. Dla każdej takiej kategorii LLM wyciągnął argumenty obu stron, oszacował ekspozycję ryzyka i pokazał mi **trade-off liczbowy** - ile podatku oznacza interpretacja konserwatywna, ile mniej restrykcyjna, jakie jest ryzyko sporu z US. **Decyzję podejmowałem ja**, nie agent. Agent zostawił uzasadnienie w `data/sources.md` w repo - gdyby kiedyś przyszła kontrola, mam udokumentowaną ścieżkę myślenia. I jeszcze jeden klucz, który warto powtórzyć każdemu, kto stoi przed pierwszą samodzielną deklaracją: **korekta PIT-38 jest możliwa do 5 lat wstecz** (do 2030 dla deklaracji za 2025). Złożenie z dobrą wiarą plus dokumentacja decyzji = bezpieczna ścieżka. Gdy nie jesteś pewien - interpretacja konserwatywna w pierwszej deklaracji zawsze działa, można potem skorygować na korzyść podatnika. LLM w tej części nie podejmuje decyzji za mnie. **Mapuje opcje, dokumentuje argumenty, czeka na input.** To dokładnie ten podział pracy, który chcę. ## "Wrzuciłeś dane finansowe do LLM?" - świadomy wybór, nie nieostrożność Wiem, że dla części czytelników już samo _"wrzuciłem dane finansowe do LLM"_ jest dyskwalifikujące. Trzy rzeczy do rozważenia. **a) Claude Code (Anthropic API) nie używa danych do treningu modeli domyślnie.** To jest inny model biznesowy niż consumer ChatGPT - i jakościowa różnica. Aktualne zasady przetwarzania zawsze warto sprawdzić w [polityce Anthropic](https://www.anthropic.com/privacy), ale baseline jest taki: API to środowisko produkcyjne dla developerów i firm, nie zbieracz danych treningowych. **b) Repo jest prywatne, lokalne.** Nie ma `git push` do GitHuba. CSV-ki są w `.gitignore` - nie wchodzą nawet do historii commitów. Bazowe dane finansowe siedzą tylko na moim dysku. Kontekst LLM-a kończy się z konwersacją. **c) Realny benchmark.** Alternatywą był księgowy z biurka, Excel na pendrive lub Twój e-PIT przeglądany w przeglądarce - w każdym z tych scenariuszy moje dane przechodzą przez czyjeś ręce, czyjąś pamięć albo czyjeś serwery. **Wybór nie jest między "bezpieczne" a "ryzykowne". Jest między różnymi rodzajami zaufania.** Świadomie wybrałem zaufanie do Anthropic plus lokalnego workflow. Ktoś inny wybierze inaczej i to OK. Dyskusja _"czy LLM jest bezpieczny dla finansów"_ traci sens, jeśli nie porównujemy go z konkretną alternatywą. ## Czego nie polecam - **Nie automatyzowałem zapłaty 172 PLN.** Przelew na mikrorachunek poszedł ręcznie. Nie ma sensu robić tego przez LLM - to jedna kwota, jeden numer rachunku, 30 sekund w aplikacji bankowej. - **Nie polecam tego setupu osobie bez programistycznego komfortu.** `/ingest`, struktura katalogów, git, `.gitignore`, edycja markdown - wymaga rozumienia narzędzi. Bez tego strata czasu na setup zje wszystkie zyski. - **LLM nie zastępuje doradcy podatkowego.** W sytuacjach niejednoznacznych - a takich jest sporo przy multi-source krypto + akcje + dywidendy - finalna decyzja musi należeć do człowieka, idealnie po konsultacji z doradcą. LLM dostarcza argumenty i mapuje ryzyko. Doradca daje rekomendację dostosowaną do Twojej sytuacji i bierze za nią odpowiedzialność. ## Jeśli czytasz to dziś - masz jeszcze ~30 godzin Jeśli jeszcze nie złożyłeś PIT-38, a masz multi-source przychody, oto minimum-viable plan na **2-3 godziny** przed deadlinem 30.04: 1. Pobierz wszystkie PIT-8C ze swoich brokerów (XTB, mBank, etc.) i raporty z giełd krypto. 2. Wejdź na Twój e-PIT - sekcja C (papiery + fundusze) jest dla większości osób auto-wypełniona. 3. Sekcję E (krypto) i G (dywidendy zagraniczne) wypełniasz ręcznie. To tutaj Twój e-PIT nie pomoże. 4. **Sprawdź swoje PIT-38 z 2024.** Czy odpowiednia pozycja zawiera niewykorzystane koszty krypto z lat ubiegłych. To może być wart **kilka-kilkadziesiąt tysięcy PLN bufor**, o którym nie pamiętasz. 5. Złóż przez profil zaufany. Zapłać mikrorachunek do 30.04. 6. **Najgorszy scenariusz: korekta PIT-38 możliwa do 2030.** Złożenie z grubsza poprawnej deklaracji w terminie jest zawsze lepsze niż brak deklaracji + czynny żal. **Po terminie**, w trzech profilach: - **Prosty PIT (1 PIT-37 z pracy)** - Twój e-PIT i tyle. Nie kombinuj. - **2-3 źródła (akcje + krypto)** - warto rozważyć 1-2 wieczory na setup workflow zbliżony do tego. - **5+ źródeł i historia strat z lat ubiegłych** - to **TWÓJ scenariusz**. Bufor kosztów krypto może być wart 5-50k PLN przeoczonego podatku rocznie. Setup się zwraca. Wartość LLM rośnie nie liniowo, ale **skokowo**, gdy struktury repo są dla niego czytelne. Bez `CLAUDE.md` i konwencji projektowych ten projekt zająłby tyle samo czasu co praca z księgowym. Z nimi - dwie godziny. To jest dokładnie ten zakres ROI, którego nie da się sprzedać generic content marketingiem - bo wymaga, żeby najpierw stała tam infrastruktura.

Masz multi-source podatki i myślisz "ja też tak chcę"?

Pomagam freelancerom i konsultantom technologicznym ustawiać AI workflow do rzeczy, które do tej pory delegowali ekspertom. Pokażę, jak taki setup mógłby wyglądać u Ciebie.

Umów bezpłatną konsultację
## Przydatne zasoby - [Twój e-PIT](https://www.podatki.gov.pl/pit/twoj-e-pit/) - automatyczne wypełnienie sekcji C dla papierów i funduszy - [Ustawa o PIT - ISAP](https://isap.sejm.gov.pl/) - art. 17 ust. 1 pkt 11, art. 22 ust. 14, 16, art. 11a (kursy), art. 30a (dywidendy) - [NBP - kursy średnie](https://nbp.pl/statystyka-i-sprawozdawczosc/kursy/) - D-1 dla transakcji walutowych - [Anthropic - Privacy & Data Usage](https://www.anthropic.com/privacy) - domyślne zasady przetwarzania w API - [Skills 2.0 - multi-agent system do zarządzania firmą](/blog/skills-2-0-multi-agent-system-zarzadzanie-firma) - kontekst, jak zbudowane są konwencje `CLAUDE.md` - [Spec-driven SEO na portfolio i Qamera AI](/blog/spec-driven-seo-portfolio-qamera-ai) - inny case study, ten sam typ workflow ## FAQ
### Czy mogę zrobić PIT-38 z Claude Code, jeśli nie jestem programistą? Krótka odpowiedź: niekoniecznie warto. Setup wymaga znajomości git, terminala, struktury katalogów, edycji markdown i `.gitignore`. Bez tego strata czasu na konfigurację zje wszystkie zyski. Bezpieczniejsza ścieżka dla osób nietechnicznych to księgowy lub Twój e-PIT z ręcznym uzupełnieniem sekcji E i G.
### Czy Anthropic używa moich danych finansowych do trenowania modeli? Domyślnie nie - Claude Code (Anthropic API) działa na innym modelu biznesowym niż consumer ChatGPT. Dane przesyłane przez API nie są używane do treningu modeli bez wyraźnej zgody klienta. To inna kategoria niż darmowe narzędzia konsumenckie. Aktualne zasady warto sprawdzić w [polityce prywatności Anthropic](https://www.anthropic.com/privacy).
### Co to jest "bufor kosztów krypto" i czemu może być wart kilkadziesiąt tysięcy PLN? Bufor to udokumentowane wydatki na zakup krypto, których jeszcze nie zamknąłeś sprzedażą na walutę tradycyjną. W polskim PIT (art. 22 ust. 14, 16) koszty czekają w buforze do roku, w którym sprzedasz krypto na fiat - wtedy obniżają podstawę opodatkowania. Nadwyżka kosztów nad przychodami w danym roku przechodzi na kolejne lata **bez ograniczenia czasowego**. Dlatego sprawdzenie zeszłorocznej deklaracji potrafi być warte kilka-kilkadziesiąt tysięcy PLN przeoczonego bufora.
### Co zrobić, jeśli czytam to po 30 kwietnia i jeszcze nie złożyłem PIT-38? Złóż jak najszybciej z czynnym żalem (art. 16 KKS) - kara za niezłożenie deklaracji rośnie z czasem, a sam czynny żal w wielu przypadkach pozwala uniknąć grzywny. Korekta PIT-38 jest możliwa do 5 lat wstecz, więc lepiej złożyć z grubsza poprawną deklarację z opóźnieniem niż w ogóle. Po fakcie warto skonsultować się z doradcą podatkowym, jeśli sytuacja jest złożona.
### Czy LLM zastępuje doradcę podatkowego przy multi-source PIT? Nie. LLM dobrze radzi sobie z klasyfikacją typów transakcji, ekstrakcją danych z PDF-ów i mechaniczną konwersją kursów NBP, ale w sytuacjach niejednoznacznej kwalifikacji prawnej finalna decyzja musi należeć do człowieka. Idealnie - po konsultacji z doradcą, który weźmie odpowiedzialność za rekomendację dostosowaną do Twojej sytuacji. LLM mapuje opcje i ryzyka. Doradca podejmuje decyzję, którą podpisuje swoim nazwiskiem.
--- # Dlaczego nie da się tego zrobić na WordPressie - spec-driven SEO na portfolio i Qamera AI Source: https://pawel.lipowczan.pl/blog/spec-driven-seo-portfolio-qamera-ai Published: 2026-04-26 ## Dwa projekty, dwa stacki, jedna pętla pracy W ciągu dwóch tygodni zoptymalizowałem SEO na dwóch radikalnie różnych projektach. **Portfolio** ([pawel.lipowczan.pl](https://pawel.lipowczan.pl)) - Vite 7 + React 19 SPA z prerenderem. Audyt znalazł 10 findings, naprawiłem pięć z nich w jedno popołudnie. Securityheaders.com przeszedł z **C na A**, Rich Results Test z **5 warnings na 0**, sitemap dostał per-URL `lastmod` zamiast jednego build timestampa dla 73 URL-i. **Qamera AI** ([qamera.ai](https://qamera.ai)) - mój SaaS do AI product photography, Next.js 16 App Router + Turborepo + Vercel + Supabase + i18n EN/PL/UK. Audyt zwrócił health score **56/100**. W pięć dni roboczych zamknąłem **dziewięć spec-driven changes** (siedem planowanych + dwa wykryte po drodze), które rozwiązały wszystkie "Critical" findings. CLS na `/marketplace/styles` spadł z **0.467 do 0.016** - 27× poprawa. Hreflang pokrycie urosło z "tylko docs" do **20 static marketing paths plus docs**. Homepage dostał trzy bloki JSON-LD (Organization + WebSite + SoftwareApplication), pricing kolejne trzy (Product × 2 + FAQPage). Centralna teza tego artykułu jest prosta i niewygodna dla części czytelników: **pełna kontrola nad SEO i GEO jest możliwa tylko przy stacku opartym na kodzie**. WordPress, Webflow i Wix dają wtyczki - nie dają nagłówka `Content-Security-Policy` z reportingiem do Sentry, nie dają `xhtml:link` na poziomie sitemapy, nie dają `requestIdleCallback` w ``, nie dają `llms.txt` generowanego z własną logiką build-time. Drugi multiplikator: dobry **AI workflow** - brainstorm → spec → execute → review → test. Sam kodowy stack bez procesu = dwa tygodnie ręcznej pracy. Sam AI workflow na zamkniętej platformie = uderzasz w sufit pluginów. Razem = godziny. Pokażę proces na obu projektach. Zobaczysz, co transferuje się 1:1, a co wymaga innych decyzji per stack. ## Dlaczego "platforma vs kod" to dziś nie debata o cenie hostingu Pięć lat temu wybór między WordPressem a własnym kodem był pragmatyczny. WordPress dawał motywy, pluginy, ekosystem, panel dla nietechnicznego klienta. Vercel z własnym frameworkiem był overkillem dla 80% projektów. W 2026 te proporcje się przesunęły. **SEO przesunęło się w stronę GEO** (Generative Engine Optimization) - ChatGPT web search, Perplexity, Claude Search i Gemini Deep Research czytają twoje strony, ale inaczej niż Googlebot. Respektują [llmstxt.org](https://llmstxt.org/) spec, ważą `author.name` i `datePublished` w JSON-LD przy wyborze źródeł, preferują "factual-definition opener" w pierwszych 150 słowach. Security headers stały się sygnałem zaufania. Core Web Vitals weryfikuje field data z CrUX, nie lab score z Lighthouse. Schema enrichment przekłada się na rich results w SERP. Co WordPress / Webflow daje w 2026: SEO plugin (Yoast, Rank Math), basic schema dla Article i Product, sitemap generowany automatycznie, redirecty, `meta description`. To około **80% potrzeb** dla typowej strony firmowej. Czego nie daje (lub daje z bardzo dużą walką): ```text - Content-Security-Policy Report-Only z reportingiem do Sentry - Permissions-Policy per-page (geolocation, camera, microphone, payment) - requestIdleCallback dla third-party scriptów zamiast async=true - xhtml:link w sitemapie (nie tylko hreflang w head) - llms.txt / llms-full.txt z własną logiką generacji - BlogPosting z mainEntityOfPage + publisher (raster logo) + ISO 8601 datetime - Per-bot reguły w robots.txt (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) - Sitemap lastmod per page-type (post: frontmatter.modified, listing: max, legal: git mtime) ``` To jest top 20% kontroli. I to dokładnie ten zakres, w którym dziś wygrywasz pozycje - w klasycznym SERP i w odpowiedziach LLM-ów. Pluginy domykają top 80%. Top 20% wymaga edycji nagłówków HTTP, struktury HTML w ``, własnego buildera artefaktów. Tego nie robisz w admin panelu - bo admin panel celowo nie odsłania tych warstw. ## Toolchain - pięć narzędzi, jedna pętla Używam tych samych pięciu narzędzi na każdym projekcie SEO. Każde z osobna jest zwykłe. Razem tworzą pętlę, która kompresuje czas o 80-90%. | Narzędzie | Rola | | ---------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | | **`claude-seo`** plugin (20+ sub-skilli) | Audyt jako pierwsza komenda - technical, GEO, schema, performance, hreflang, AI discoverability | | **OPSX / OpenSpec** | Spec-driven workflow - `proposal.md` → `design.md` → `specs/` → `tasks.md` przed kodem | | **Lighthouse MCP** | Lab CWV i LCP opportunities z poziomu agenta, bez przełączania kontekstu do PageSpeed | | **Rich Results Test + securityheaders.com + Sentry CSP Reports** | Weryfikacja na każdym kroku | | **Git worktrees** | Równoległa praca nad niezależnymi changes (gdy projekt na to pozwala) | OPSX opisałem szerzej w [osobnym artykule o strukturyzowanym podejściu do AI workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai). `claude-seo` to przykład **specjalizowanego skilla** w sensie z [posta o Skills 2.0](/blog/skills-2-0-multi-agent-system-zarzadzanie-firma) - domain-specific knowledge plus checklist plus pre-built sub-skille. Pętla pracy wygląda tak samo na każdym projekcie: ```text ┌─────────┐ ┌──────────┐ ┌────────┐ ┌───────┐ │ Audit │ → │ Proposal │ → │ Design │ → │ Specs │ └─────────┘ └──────────┘ └────────┘ └───────┘ │ ┌─────────┐ ┌────────┐ ┌──────────┐ ┌───▼───┐ │ Archive │ ← │ Verify │ ← │ Implement│ ← │ Tasks │ └─────────┘ └────────┘ └──────────┘ └───────┘ ``` Typowa sekwencja komend: ```bash # Faza 1: audyt /claude-seo:seo # Faza 2: change proposal /opsx:new "seo-improvements" /opsx:ff # fast-forward - wszystkie planning artifacts naraz # Faza 3: implementacja /opsx:apply # Faza 4: weryfikacja /opsx:verify # manualnie: securityheaders.com, Rich Results Test, Lighthouse MCP # Faza 5: archiwizacja /opsx:archive ``` Każde z tych narzędzi działa **dlatego, że substratem jest kod**. Audyt może czytać dowolny plik źródłowy i raw output ``. OPSX może edytować `vite.config.js`, `next.config.ts`, `vercel.json`, `robots.txt` w jednej sesji. Weryfikatory dostają pełny output bez sandboxa. To nie jest przypadek - to konsekwencja architektury. ## Audyt - co znajduje claude-seo na dwóch radikalnie różnych projektach Pierwsza obserwacja, która mnie zaskoczyła: **claude-seo zwraca ten sam zestaw kategorii znalezisk niezależnie od stacku**. Różny jest tylko sposób ich naprawy. | Projekt | Stack | Findings | Stan początkowy | | --------- | ------------------------------- | --------------------------------------------------- | ------------------------------------------ | | Portfolio | Vite 7 + React 19 + Vercel | 10 (4 perf, 2 schema, 2 security, 1 sitemap, 1 GEO) | securityheaders C, Rich Results 5 warnings | | Qamera AI | Next.js 16 + Turborepo + Vercel | 7 + 2 wykryte podczas | health score 56/100 | Wspólne kategorie znalezisk, które transferują się 1:1 między stackami: - **Brak lub niekompletny `llms.txt`** - biggest miss na GEO readiness w obu projektach - **Schema enrichment** - brakujący `publisher`, `dateModified`, `mainEntityOfPage`, ISO 8601 datetime - **Hreflang tylko head-level**, brak `xhtml:link` w sitemap dla klasteryzacji wariantów językowych - **Security headers** - deprecated `X-XSS-Protection`, brak `Permissions-Policy`, brak HSTS preload, brak CSP - **AI bot allowlist** to wildcard - co dla `Google-Extended` i `GPTBot` oznacza "brak sygnału", nie "allow" Stack-specific findings, które wymagają innych decyzji: - **Portfolio:** `clickrank.ai` synchroniczny w `` blokuje parser przed First Paint, sitemap z 73 URL-ami i jedną datą `lastmod`, `articleBody: post.excerpt` semantycznie błędne w `BlogPosting` - **Qamera:** CLS 0.467 na `/marketplace/styles` przez client-side fetch z Airtable bez zarezerwowanych wymiarów kart, hardcoded EN strings w `root-metadata.ts` (PL/UK użytkownicy dostawali angielski OG na każdej marketingowej), Merchant Listings false-positive na pricing Wspólny zestaw znalezisk to **pierwszy dowód, że proces jest transferowalny**. Zna mnie kilka pluginów SEO i każdy z nich na obu projektach zwróciłby fundamentalnie różne raporty, bo każdy jest zwiazany z konkretną platformą. Audyt agenta na neutralnym substracie - kodzie - zwraca uniwersalny obraz. ## Od audytu do change proposal - kiedy bundlować, kiedy splitować Dwa projekty, dwie różne strategie packagingu zmian. **Portfolio** dostało jeden change `seo-improvements` z pięcioma filarami w jednym PR-ze ([#2](https://github.com/plipowczan/portfolio/pull/2)). Single-maintainer, brak ryzyka konfliktu plików, łatwiejszy review całości - bo zmiany są logicznie związane tematycznie. **Qamera** dostała dziewięć osobnych changes, osiem PR-ów (#75/76/77/82/92/93/94/96), pracowanych równolegle na worktree'ach git. Multi-developer, monorepo, disjoint file sets. Worktrees pozwoliły każdej zmianie mieć własne `node_modules` i własny port dev servera - zero konfliktu state. Kryterium decyzyjne, którego używam: | Czynnik | One-PR (portfolio) | Multi-PR (Qamera) | | ------------------------- | ------------------ | ------------------------------ | | Liczba maintainerów | 1 | 2+ | | Ryzyko konfliktu plików | niskie | wysokie | | Cykl review | self-review | code review przez wspólnika | | Rozkład czasowy | jedno popołudnie | 5 dni roboczych | | Rollback granularity | całość lub nic | per-feature | | Dev environment isolation | nie potrzebna | worktree + osobne node_modules | Wspólne dla obu: każda zmiana = OPSX `proposal.md` + `design.md` + `specs/` + `tasks.md` **przed** kodem. To nie biurokracja. To **feedback loop dla AI**: review specu kosztuje minuty, review 200 linii wygenerowanego kodu w niewłaściwym miejscu kosztuje godziny. Spec-driven daje ci punkt weta, zanim zapłacisz koszt implementacji. ```text ## Tasks - seo-improvements - [x] Move clickrank inline to requestIdleCallback (+ setTimeout fallback) - [x] Generate dedicated raster logo (600×60 PNG via sharp) - [x] Add publisher / dateModified / mainEntityOfPage to BlogPosting - [x] Drop articleBody: excerpt (semantically wrong) - [x] Build llms.txt / llms-full.txt generator (scripts/generate-llms-txt.js) - [x] Replace X-XSS-Protection with Permissions-Policy + HSTS preload - [x] Configure CSP Report-Only → Sentry Security Reports bucket - [x] Per-page-type lastmod (post: frontmatter.modified, listing: max, legal: git mtime) ``` Każdy task w `tasks.md` to jedna jednostka pracy o znanym scope i znanym sposobie weryfikacji. Po implementacji checkbox jest dowodem, że task został wykonany - nie deklaracją. ## Co transferuje się 1:1 (i dlaczego to argument za kodem) Cztery rzeczy, które zaimplementowałem w identycznym wzorcu na portfolio i w Qamerze. Każda byłaby trudna lub niemożliwa na zamkniętej platformie. ### A. `llms.txt` jako własny artefakt build-time Spec [llmstxt.org](https://llmstxt.org/) istnieje od 2024 roku (Answer.AI / Jeremy Howard). W 2026 ChatGPT web search, Perplexity, Claude Search i Gemini Deep Research go respektują. Plik to skrócony index treści dla LLM, z opcjonalnym `llms-full.txt` zawierającym pełną treść do single-token ingest. Na portfolio mam `scripts/generate-llms-txt.js` uruchamiany w `build:prerender`. Czyta `src/content/blog/*.md` (PL + EN) przez `gray-matter`, plus `src/data/projects.js`. Generuje `public/llms.txt` (~16 KB index) i `public/llms-full.txt` (~800 KB pełna treść z separatorem `\n\n---\n\n`). ```text # Pawel Lipowczan > Architekt oprogramowania i doradca ds. technologii... ## Blog (PL) - [Tytuł](url): jednozdaniowy opis ... ## Blog (EN) - [Title](url): one-line description ... ## Kontakt - email: ... ``` W Qamerze bliźniaczy skrypt w workspace `apps/web` generuje `llms.txt` z marketingowych stron, blog postów i public docs. Logika jest inna (źródła danych, struktura sekcji), ale wzorzec - build-time generator zgodny ze spec - identyczny. **Tego nie zrobisz w panelu:** plugin WordPressa może wyplunąć statyczny `llms.txt`, ale nie zaintegrujesz go z własnym CMS-em na własnych warunkach (kolejność sekcji per język, fallback dla brakujących `description`, paginacja przy 100+ artykułach). ### B. Schema enrichment - `articleBody: excerpt` to semantyczny błąd Mój stary `BlogPosting` miał sześć pól. Rich Results Test pokazywał pięć non-critical warnings. Po enrichmencie - jedenaście pól, zero warnings. ```json { "@type": "BlogPosting", "headline": "...", "description": "pierwsze 300 znaków contentu lub frontmatter.description", "author": { "@type": "Person", "name": "Pawel Lipowczan", "url": "https://pawel.lipowczan.pl" }, "datePublished": "2026-01-15T00:00:00Z", "dateModified": "2026-04-21T00:00:00Z", "image": "...", "url": "...", "mainEntityOfPage": { "@type": "WebPage", "@id": "..." }, "publisher": { "@type": "Organization", "name": "Pawel Lipowczan", "logo": { "@type": "ImageObject", "url": "https://pawel.lipowczan.pl/logo-schema.png" } } } ``` Trzy nieoczywiste detale: `articleBody: post.excerpt` jest **semantycznie błędne** (spec wymaga pełnej treści, nie skrótu) - wyciąłem to pole zupełnie. `publisher.logo` musi być **rasterem** (PNG 600×60), nie SVG. ISO 8601 z `Z` lub offsetem, nie `2026-01-15` bez strefy. W Qamerze identyczny zestaw zmian dotknął `Article` na `/blog`, `Service` na `/offer/*` i `Product` na `/pricing`. **Tego nie zrobisz w panelu:** SEO pluginy ustawiają top sześć pól. `mainEntityOfPage`, `publisher.logo` jako oddzielny raster, ISO datetime, `description` z fallback do pierwszego akapitu - to ręczna robota w generatorze schema. ### C. Hreflang na poziomie sitemapy, nie tylko `` Next.js `Metadata.alternates.languages` to head-level signal. Google preferuje **sitemap-level `xhtml:link`** dla klasteryzacji wariantów językowych. W Qamerze rozwiązaliśmy to wspólnym helperem `buildLanguageAlternates(pathname)` używanym z dwóch miejsc - `sitemap.ts` i każdego `generateMetadata` per page. ```xml https://qamera.ai/pricing ``` Drift-guard test w CI failuje, gdy ktoś doda ścieżkę do sitemap, a nie doda `alternates` do `page.tsx`. To zabezpieczenie przed silent regression - bardzo łatwo dodać nowy landing i zapomnieć o jego wariancie językowym. **Tego nie zrobisz w panelu:** Yoast generuje hreflang w head. Sitemap-level wymaga edycji generatora sitemapy plus drift-guard - czyli kodu w CI. ### D. AI bot allowlist - named rules zamiast wildcard `robots.txt` z osobnymi blokami dla `GPTBot`, `OAI-SearchBot`, `ClaudeBot`, `PerplexityBot`, `Google-Extended` i `CCBot`. Wildcard = "brak sygnału" - bot interpretuje to konserwatywnie. Named allow = "explicit yes" - bot wie, że może crawl-ować i indeksować dla swojego pipeline. **Tego nie zrobisz w panelu:** WordPress pisze do `robots.txt` przez plugin, ale per-bot reguły wymagają edycji pliku fizycznego - czyli dostępu do filesystem, którego nie masz w typowym shared hostingu. ## Co jest stack-specific (i czego nauczył mnie każdy projekt osobno) ### Portfolio - `async=true` na inline script to mit Ten finding zaskoczył mnie najbardziej, bo wszyscy się mylą. Skrypt `clickrank.ai` w `` wyglądał tak: ```html ``` Pułapka: `async=true` dotyczy **ściągania** skryptu, ale sam inline kod, który go tworzy, wykonuje się **synchronicznie podczas parsowania HTML**. Dodaje microtask do event loopa, zanim browser wyrenderuje cokolwiek. Fix - `requestIdleCallback` plus fallback dla Safari 16.3 i starszych: ```html ``` Weryfikacja po deploy na prod: `performance.getEntriesByType('resource').filter(r => r.name.match(/clickrank/))` → `startTime: 101.6ms`. Browser zgłosił idle po ~100ms i dopiero wtedy odpalił callback. Lighthouse lab score variance pozostała duża (post: prod 38 → preview 61 → drugi run 43) - **lab score ≠ field data**. Prawdziwa weryfikacja to CrUX z Google Search Console po 2-4 tygodniach. ### Qamera - CLS 0.467 → 0.016 przez SSR initial grid `/marketplace/styles` pokazywał karty stylów ładowane client-side z Airtable. Bez zarezerwowanych wymiarów grid layout shiftował się 4× ponad próg failing (0.467 vs cel ≤ 0.1). Trzy opcje fixu - SSR initial grid, reserved card dimensions, combined. Wybraliśmy **SSR**: bonus dla GEO (non-JS crawlers widzą content) plus eliminacja CLS u źródła. Rezultat: **CLS 0.016** (27× poprawa), LCP 2.4s → 1.6s. Twist post-deploy: PageSpeed Insights pokazał LCP 14.4s (cold Vercel function), Lighthouse MCP równolegle 1.6s (warm). **Jedna metryka z PSI to sampling.** Zawsze re-run lub weryfikuj lokalnie. ## Bug, którego audyt nie szukał - i dlaczego to argument za regularnymi audytami Audyt SEO portfolio wyrzucił finding, którego się nie spodziewałem: hreflang alternatywy dla posta `llm-knowledge-base-brain-karpathy` wskazywały na `/en/blog/`, który zwraca "Post not found". Root cause: post był **PL-only** (brak EN wersji), ale jego frontmatter miał `alternateSlug: llm-knowledge-base-brain-karpathy` - wskazujący sam na siebie. Efekt wcześniejszej iteracji blog-article-writer skilla, który **autouzupełnił pole bez walidacji**. Łańcuch zdarzeń: 1. User na PL poście klika przełącznik języka 2. `getAlternatePost(currentSlug)` zwraca... ten sam PL post 3. LanguageSwitcher buduje `/en/blog/` i nawiguje 4. `BlogPostPage` filtruje `getPostsByLang("en")` → brak match → "Post not found" 5. Sitemap dziedziczy ten bug jako bad hreflang, propaguje do Google Trzy-poziomowa naprawa: **data fix** (usunięcie pola), **code defense** (`getAlternatePost` odrzuca self-reference i same-lang candidates), **process fix** (reguła w `.claude/rules/data-storage/` plus update walidatora w blog-article-writer skillu). Meta-lekcja jest mocna: audyt SEO uruchamia bug-i, **które nie były jego celem**. Nigdy bym nie znalazł tego bez claude-seo. To argument za regularnym audytem nawet na małym projekcie. Drugi meta-poziom - bug został **wprowadzony przez AI workflow** (blog-article-writer skill), naprawiony przez **inny AI workflow** (audyt + spec-driven fix + reguła w skillu). To pętla samokorygująca pod warunkiem, że jest proces. Bez procesu - bug żyłby tygodniami. Pokrewny wątek o tym, jak agent pilnuje swoich własnych standardów, opisałem szerzej w [poście o LLM Wiki Karpathy'ego](/blog/llm-knowledge-base-brain-karpathy) i [Second Brain z Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills). ## Kompresja czasu jest multiplikatywna, nie addytywna Liczbowo: - **Portfolio:** audyt 15 min + 4h implementacji + 30 min weryfikacji = **5h** dla 5 zmian - **Qamera:** **5 dni roboczych** dla 9 zmian (siedmiu planowanych + dwóch wykrytych po drodze) - **Drugi projekt = ~30% czasu pierwszego** dzięki transferowi wzorców (llms.txt, schema enrichment, hreflang sitemap, AI bot allowlist) Multiplikator: **kodowy stack × dobry AI workflow = godziny**. Każdy z osobna nie wystarczy. Sam kodowy stack bez procesu = dwa tygodnie ręcznej pracy z forami i Stack Overflow. Sam AI workflow na zamkniętej platformie = uderzasz w sufit pluginów po pół godziny. Razem dają kompresję o 80-90%. To nie jest addytywne - to mnożenie. Argument, dla którego coraz częściej wybieram kodowe rozwiązania, opisałem też w [przewodniku po vibe codingu](/blog/vibe-coding-przewodnik). Sześć takeaways z obu projektów: 1. **Kodowy stack daje top 20% kontroli, której pluginy nie dają** - i to ten zakres dziś wygrywa pozycje 2. **Spec-driven jako feedback loop dla AI** - review specu kosztuje minuty, review 200 linii kodu kosztuje godziny 3. **Transferowalne 1:1 między stackami:** llms.txt, schema enrichment, hreflang sitemap-level, AI bot allowlist 4. **Stack-specific:** każdy framework ma swoje pułapki performance i własne API metadata - tu zaoszczędzisz najmniej 5. **Audyt znajduje bug-i poza swoim scopem** - `alternateSlug === slug` nigdy nie był na liście, znalazłem przez claude-seo 6. **Drugi projekt = 30% czasu pierwszego** - pod warunkiem dokumentacji wzorców Jeśli zostajesz na WordPressie - ten artykuł nie zmienia twojego życia. Jeśli rozważasz przejście na własny stack, to argument, którego potrzebowałeś.

Potrzebujesz audytu SEO + GEO na własnym stacku?

Robię to samo na projektach klientów - od audytu przez spec-driven changes po post-deploy verification. Omówimy twój stack i realny scope w 30 minut.

Umów bezpłatną konsultację
## Przydatne zasoby - [llmstxt.org](https://llmstxt.org/) - spec llms.txt - [securityheaders.com](https://securityheaders.com/) - skaner nagłówków bezpieczeństwa - [Rich Results Test](https://search.google.com/test/rich-results) - walidator structured data Google - [PageSpeed Insights](https://pagespeed.web.dev/) - Core Web Vitals lab + field data - [MDN - requestIdleCallback](https://developer.mozilla.org/en-US/docs/Web/API/Window/requestIdleCallback) - [MDN - Content-Security-Policy](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Content-Security-Policy) - [Sentry Security Reports](https://docs.sentry.io/product/security-policy-reporting/) - CSP reports via Sentry - [Google Search Central - Article structured data](https://developers.google.com/search/docs/appearance/structured-data/article) - [Qamera AI](https://qamera.ai) - projekt opisany w case study ## FAQ
### Czy każda strona oparta na kodzie jest automatycznie lepsza pod SEO niż WordPress? Nie - kodowy stack daje **kontrolę**, nie wynik. Bez procesu (audyt → spec → execute → weryfikacja) skończysz z gorszą stroną niż dobrze skonfigurowany WordPress z Yoastem. Argument tego artykułu jest taki: kodowy stack **pozwala** zoptymalizować top 20% (CSP, llms.txt, sitemap-level hreflang, schema enrichment), których platformy nie odsłaniają. Czy to wykorzystasz, zależy od twojego workflow.
### Co to jest spec-driven development w kontekście SEO? Spec-driven oznacza, że każda zmiana zaczyna się od artefaktów: `proposal.md` (co i dlaczego), `design.md` (jak), `specs/` (kontrakty), `tasks.md` (lista kroków) - **przed** napisaniem kodu. W SEO sprawdza się szczególnie, bo zmiany dotykają wielu warstw (HTTP headers, HTML head, structured data, sitemap), a brak specu = AI generuje 200 linii kodu w niewłaściwym miejscu. Używam OpenSpec / OPSX workflow - szczegóły w [osobnym artykule](/blog/opsx-workflow-strukturyzowana-praca-z-ai).
### Czy llms.txt ma sens w 2026, jeśli moja strona nie jest tutorialem AI? Tak, ale ROI jest niższy. `llms.txt` najmocniej działa dla treści, które LLM-y cytują (tutoriale, dokumentacja, case studies). Dla e-commerce lub portfolio impact jest mniejszy, ale wciąż dodatni - koszt to 100-200 linii skryptu Node, korzyść to obecność w grounding ChatGPT, Perplexity i Claude Search. Plik `llms-full.txt` przy 100+ artykułach robi się ciężki - wtedy paginacja albo `top-articles-only`.
### Jak wybrać między jednym dużym PR-em a wieloma małymi przy zmianach SEO? Single-PR ma sens przy single-maintainerze i tematycznie spójnych zmianach (jak portfolio: 5 filarów SEO w jednym PR-ze, 4h pracy). Multi-PR jest konieczny przy wielu maintainerach, monorepo i równoległej pracy (jak Qamera: 9 zmian, 8 PR-ów, 5 dni). Kryterium: czy zmiany dotykają tych samych plików (konflikt = split) i czy review całości jest realny w jednym przejściu (>500 linii diff = split).
### Czy AI workflow zastępuje code review przy zmianach SEO? Nie - uzupełnia. W Qamerze Copilot review na PR złapał trzy trafne issues (placeholder Sentry DSN, brak preview env var, unused import), których spec-driven workflow nie złapał. AI workflow przyspiesza generację kodu zgodnego ze specem, ale **drugi pair of eyes** (człowiek lub AI reviewer) wciąż łapie różnicę między "kod robi to, co spec mówi" a "kod robi to, co spec mówi, w sposób bezpieczny dla produkcji".
--- # Jak LLM Wiki Karpathy'ego pomogła mi uporządkować moją bazę wiedzy Source: https://pawel.lipowczan.pl/blog/llm-knowledge-base-brain-karpathy Published: 2026-04-12 ## Karpathy opisał framework. Ja miałem żywy system do uporządkowania Andrej Karpathy w kwietniu 2026 opublikował na X [wątek o "LLM Wiki"](https://x.com/karpathy/status/2039805659525644595) - koncepcji, w której LLM buduje i utrzymuje persistent wiki z Twoich źródeł. Zamiast klasycznego RAG, który szuka fragmentów na żądanie, LLM aktywnie zarządza bazą wiedzy: tworzy notatki, aktualizuje cross-referencje, flaguje sprzeczności. Czytam ten wątek i mam déjà vu. Nie dlatego, że zbudowałem to samo - ale dlatego, że od 2022 roku organicznie ewoluowałem w tym kierunku. Mój vault w Obsidian zaczynał jako klasyczny zbiór ręcznych notatek - kilkadziesiąt plików Markdown, ręcznie linkowanych, rosnących bez jasnej struktury. Z czasem **Claude Code** przejął coraz więcej pracy: najpierw proste formatowanie, potem indeksowanie, wreszcie pełne zarządzanie strukturą i standardami. W pewnym momencie zdałem sobie sprawę, że agent robi więcej maintenance niż ja. Repozytorium Karpathy'ego - [LLM Wiki gist](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) - nie było dla mnie odkryciem nowego świata. Było **katalizatorem porządkowania**. Karpathy dał nazwę warstwom, które u mnie istniały chaotycznie. Pomógł mi sformalizować separation między raw sources a przetworzoną wiki. Dopisać safety rules, które wcześniej miałem tylko w głowie. Opisał navigation protocol, który u mnie działał intuicyjnie, ale nigdy nie był udokumentowany. > _"Boring part of maintaining a knowledge base isn't reading or thinking - it's bookkeeping."_ - Karpathy To jest historia ewolucji od ręcznego PKM do agentowego systemu wiedzy. Z Karpathym jako katalizatorem, który pomógł mi uporządkować to, co budowałem od lat. I z jednym centralnym konceptem, który zmienił wszystko: **progressive disclosure** - organizacja notatek tak, żeby agent AI mógł je znaleźć bez zaśmiecania okna kontekstowego. W tym artykule pokażę tę ewolucję: od chaosu ręcznych notatek, przez stopniowe przekazywanie maintenance agentowi, do systemu z 185 notatkami, trzema indeksami nawigacyjnymi i czterema workflow. Pokażę konkretną architekturę, konkretne procesy i konkretne liczby - nie teorię. ## Problem z klasycznym PKM Każdy, kto próbował utrzymać bazę wiedzy przez dłużej niż kilka miesięcy, zna ten wzorzec: na początku jest entuzjazm, potem rośnie chaos, na końcu porzucasz system. Dlaczego klasyczne podejścia nie skalują się: - **Maintenance burden rośnie szybciej niż wartość.** Każda nowa notatka to potencjalne linki do dziesiątek istniejących. Przy 50 notatkach to jeszcze ogarniam. Przy 185 - nie ma szans. Koszt utrzymania porządku rośnie wykładniczo, a wartość z każdej dodanej notatki - liniowo. - **Cross-referencje są zawsze niepełne.** Piszę notatkę o context engineering i zapominam, że trzy miesiące temu zapisałem coś powiązanego w zupełnie innym folderze. Link nigdy nie powstaje. A bez tego linku - wiedza jest rozproszona, jakby nie istniała. - **Wikis giną z powodu maintenance, nie braku wartości.** Nikt nie porzuca bazy wiedzy bo jest bezwartościowa. Porzuca ją, bo utrzymanie staje się chore. Widziałem to nie raz - w firmach, w open-source, we własnych projektach. - **Zettelkasten, BASB, Digital Garden - piękne filozofie.** Building a Second Brain Tiago Forte, Zettelkasten Luhmanna, Digital Garden - każda z tych metodologii ma sens na papierze. Ale wszystkie zakładają, że TY jesteś bottleneckiem maintenance. I mają rację - dlatego systemy umierają. Nie dlatego że ludzie są leniwi. Dlatego że bookkeeping to nie jest praca, na którą masz czas obok pisania kodu, prowadzenia firmy i czytania nowych rzeczy. Karpathy ujął to precyzyjnie: nudna część utrzymania knowledge base to nie czytanie ani myślenie. To bookkeeping - aktualizowanie cross-referencji, pilnowanie aktualności podsumowań, notowanie sprzeczności. **LLMs don't get bored.** I to jest kluczowa obserwacja. Nie chodzi o to, że AI jest mądrzejszy w organizowaniu wiedzy. Chodzi o to, że AI nie ma problemu z powtarzalnymi, nudnymi zadaniami, które ludzie odkładają na "kiedyś". Przez trzy lata z Obsidianem przeszedłem przez każdy z tych etapów. Zaczynałem z entuzjazmem, potem pojawiły się duplikaty i osierocone notatki, wreszcie doszedłem do punktu gdzie szukanie czegokolwiek zajmowało więcej czasu niż pisanie od nowa. Mój vault przetrwał, bo w pewnym momencie przestałem walczyć z maintenance ręcznie - i oddałem go agentowi. ## Architektura - jak to wygląda po uporządkowaniu ### Stack techniczny System opiera się na pięciu elementach: - **[Obsidian](https://obsidian.md/)** - edytor i frontend (lokalny, Git-based) - **[Quartz 4](https://quartz.jzhao.xyz/)** - Static Site Generator → GitHub Pages - **GitHub Actions** - automatyczny deploy na push do brancha `v4` - **Claude Code** - agent utrzymujący vault - **URL publiczny:** [brain.lipowczan.pl](https://brain.lipowczan.pl) Obsidian to interfejs dla mnie - tu czytam, przeglądam graph view, robię ad hoc notatki. Quartz 4 to interfejs dla świata - statyczna strona na GitHub Pages, dostępna pod [brain.lipowczan.pl](https://brain.lipowczan.pl). Claude Code to silnik, który pilnuje porządku między jednym a drugim. Całość jest Git repo - każda zmiana tracked, każdy ingest to commit, historia jest pełna i revertable. ### Trzy warstwy - framework Karpathy'ego, który pomógł mi uporządkować Te warstwy istniały u mnie organicznie od dawna. Miałem inbox na surowe materiały, przetworzony content i jakąś formę konfiguracji. Ale Karpathy dał im nazwy i formalną strukturę - i to pomogło mi je wyostrzyć. | Warstwa | Karpathy (LLM Wiki) | Moja implementacja | | ----------- | ----------------------- | -------------------------------------------------- | | Raw sources | Immutable drop zone | `/content/_raw/inbox/` - drop zone, nie w buildzie | | Wiki | LLM-generated .md files | `/content//` - budowane i publikowane | | Schema | Config document | `CLAUDE.md` - 300+ linii konfiguracji agenta | Separacja raw sources od wiki to kluczowa zmiana, którą sformalizowałem po przeczytaniu Karpathy'ego. Wcześniej surowe materiały i przetworzone notatki mieszały się w jednym katalogu - czasem agent modyfikował źródło zamiast tworzyć nową notatkę. Teraz inbox jest immutable drop zone - agent przetwarza, ale oryginał zostaje nietknięty w `_raw/processed/`. Zawsze mogę wrócić do źródła i porównać z tym co agent z niego zrobił. To prosty safety mechanism, ale daje spokój ducha - wiem że nic nie jest tracone. ### Progressive disclosure - dlaczego notatki są ułożone DLA agenta To centralny koncept całego systemu. Notatki nie są ułożone żeby ładnie wyglądały w Obsidian graph view. Są ułożone żeby **agent AI mógł je szybko znaleźć BEZ zaśmiecania okna kontekstowego** niepotrzebnymi informacjami. Trzy indeksy nawigacyjne - od ogółu do szczegółu: - **`vault-map.md`** (~80 linii) - bird's-eye view całego vault. Agent czyta ZAWSZE jako pierwszy. Wystarczy żeby zrozumieć strukturę i zdecydować gdzie szukać dalej. - **`catalog.md`** (~650 linii) - jedna linia per notatka z tytułem, kategorią i krótkim opisem. Agent czyta gdy potrzebuje znaleźć konkretną notatkę. - **`graph.md`** - wikilink graph (outgoing + incoming edges). Agent czyta gdy potrzebuje kontekstu powiązań między notatkami. ```text _indexes/ ├── vault-map.md ← zawsze pierwszy (~80 linii) ├── catalog.md ← 1 linia / notatka (~650 linii) └── graph.md ← wikilink edges ``` Agent nawiguje jak człowiek ze spisem treści - nie grep-uje 185 plików. Czyta vault-map (80 linii), decyduje który temat jest relevant, sięga do catalog po konkretne notatki, a jeśli potrzebuje kontekstu powiązań - otwiera graph. Dzięki temu okno kontekstowe zostaje czyste, a odpowiedzi są trafniejsze. Porównaj to z naiwnym podejściem: "wrzuć wszystkie pliki do context window i pytaj". Przy 185 notatkach po ~500 słów to ~92 500 tokenów samych notatek. Większość modeli albo tego nie pomieści, albo "zgubi się" w połowie - zacznie odpowiadać na podstawie fragmentów, które akurat trafiły blisko zapytania w embedding space, ignorując resztę. Progressive disclosure rozwiązuje ten problem elegancko: agent czyta tyle ile potrzebuje, w kolejności od ogółu do szczegółu. To jest **progressive disclosure** w praktyce. Minimalizacja zużycia context window to nie optymalizacja - to fundament, bez którego agent gubi się w szumie. Bez nawigacyjnych indeksów masz dwa wyjścia: albo dajesz agentowi cały vault (za dużo kontekstu), albo sam wskazujesz które pliki czytać (wraca ręczna praca). Indeksy eliminują oba problemy. ## CLAUDE.md - "schema" jako serce systemu Karpathy nazywa to **"schema document"** - plik definiujący jak agent powinien traktować bazę wiedzy. W moim przypadku to `CLAUDE.md` - 300+ linii konfiguracji wstrzykiwanej do system promptu Claude Code przy każdej sesji. Co zawiera moje CLAUDE.md - i dlaczego każda sekcja istnieje: - **Struktura katalogów** i przeznaczenie każdego folderu - żeby agent wiedział gdzie tworzyć notatki i czego nie ruszać - **Navigation protocol** - progressive disclosure, kolejność czytania indeksów, kiedy sięgać po pełne pliki - żeby agent nie zaśmiecał sobie context window - **Writing style guidelines** - styl pisania notatek, mix PL/EN, emoji w nagłówkach, formatowanie - żeby wszystkie notatki wyglądały spójnie niezależnie od sesji - **Frontmatter schema** - YAML metadane każdej notatki (wymagane pola, typy, walidacja) - żeby agent nie tworzył notatek z niekompletnym metadanymi - **Workflow definitions** - INGEST, COMPILE, LINT, Q&A, ENHANCE z konkretnymi krokami - żeby każdy workflow był powtarzalny i deterministyczny - **Safety rules** - czego agent nigdy nie powinien modyfikować ani usuwać (np. indeksy tylko aktualizować, nie przebudowywać od zera; nie zmieniać frontmatter istniejących notatek bez potwierdzenia) ```yaml # Fragment frontmatter schema z CLAUDE.md required_fields: - title # Tytuł notatki - category # Kategoria tematyczna (AI, BUSINESS, CODE...) - tags # Lista tagów - summary # 2-3 zdania podsumowania - created # Data utworzenia (YYYY-MM-DD) - updated # Data ostatniej aktualizacji - source # Skąd pochodzi wiedza (URL, książka, doświadczenie) ``` **Kluczowy insight:** CLAUDE.md nie jest generowane przez LLM. Jest pisane ręcznie i ewoluuje z doświadczenia. Badanie ETH Zurich (2026) pokazało, że LLM-generated agentfiles _pogarszały_ performance agentów przy **20%+ wyższym koszcie**. Powód jest prosty - LLM generuje verbose, redundantne instrukcje, które zaśmiecają context window. Ręcznie pisane instrukcje są precyzyjne i universally applicable. To kontraintuicyjne. Wydaje się, że skoro LLM pisze lepszy tekst niż większość ludzi, to powinien pisać sobie konfigurację. W praktyce - LLM generuje "na wszelki wypadek" instrukcje, powtarza się, dodaje edge case'y które nigdy nie zachodzą. Efekt: 800 linii zamiast 300, context window zaśmiecone, agent wolniejszy i mniej precyzyjny. > _"Kluczowy insight: większość failures agentów to nie problem modelu, lecz konfiguracji."_ Less is more. Każda linia w CLAUDE.md musi mieć uzasadnienie. Zaczynam od prostego, dodaję tylko co faktycznie potrzebuję, usuwam co się nie sprawdza. To żywy dokument - nie spec napisany raz i zapomniany. Moje CLAUDE.md przeszło przez kilkanaście iteracji - niektóre reguły dodałem po tym jak agent usunął ważne metadane, inne po tym jak zignorował istniejące powiązania. Każda reguła to lekcja z konkretnego failure. ## Workflows - co agent umie zrobić ### INGEST - od pliku do wiedzy w 5 minut To najczęstszy workflow - i ten, który najlepiej pokazuje wartość systemu. Kiedyś przetworzenie jednego artykułu na notatkę zajmowało mi 20-30 minut. Przy 10 źródłach - cały wieczór. Z agentem? 5 minut + review. Jak wygląda INGEST krok po kroku: 1. **Drop** - wrzucam plik do `_raw/inbox/` - artykuł, PDF, transkrypt, notatka z Web Clippera, cokolwiek w formacie tekstowym 2. **Trigger** - piszę `ingest` w Claude Code 3. **Navigate** - agent czyta `vault-map.md`, sprawdza `catalog.md` pod kątem nakładania się z istniejącymi notatkami. Jeśli temat już istnieje - proponuje aktualizację zamiast duplikatu 4. **Create** - tworzy notatkę z template'u, wypełnia frontmatter (kategoria, tagi, summary, source, data) 5. **Link** - dodaje wikilinki do powiązanych notatek - i aktualizuje te notatki żeby linkowały z powrotem (bidirectional linking) 6. **Archive** - przenosi source do `_raw/processed/YYYY-MM-DD_nazwa` - oryginał zostaje, ale poza inboxem 7. **Index** - aktualizuje wszystkie trzy indeksy (vault-map, catalog, graph) Jeden source może dotknąć **10-15 plików** w jednym passie. Notatka + aktualizacje powiązanych notatek + trzy indeksy. Ręcznie - 2 godziny przy dbałości o cross-referencje. Z agentem - 5 minut + mój review, który sprawdza czy agent poprawnie zrozumiał materiał i umieścił go we właściwym kontekście. Review to kluczowy element. Nie akceptuję wyniku INGEST na ślepo. Sprawdzam: czy kategoria jest poprawna? Czy tagi mają sens? Czy cross-referencje wskazują na faktycznie powiązane notatki, a nie na luźne skojarzenia? Czy summary oddaje istotę źródła? To zajmuje 2-3 minuty, ale daje pewność, że vault utrzymuje jakość. Agent robi 95% pracy - ja dbam o te kluczowe 5%, które wymagają ludzkiego judgmentu. ### COMPILE - syntetyzuj z wielu źródeł Agent łączy wiele notatek w skompilowany artykuł. Przykład: notatka o LLM Knowledge Bases to skompilowany artykuł z wątku Karpathy'ego, moich notatek o RAG i kontekstu z vault. Agent czyta źródła, identyfikuje wspólne wątki, syntetyzuje i tworzy `compiled-note` z pełnymi cytatami. Jak to wygląda w praktyce: mówię "compile everything I have about context engineering into a single note". Agent czyta catalog, identyfikuje 7 powiązanych notatek, czyta je, wyciąga kluczowe koncepcje i buduje spójny artykuł z sekcjami i cytatami. Rezultat: jedna comprehensive notatka zamiast 7 rozproszonych fragmentów, z wyraźnym attribution do źródeł. To workflow, którego ręcznie nigdy bym nie robił regularnie - bo wymaga przeczytania i zestawienia wielu dokumentów naraz. Przy 7 notatkach po 500 słów to 3500 słów do przeczytania, zrozumienia i zsyntezowania. Agent robi to w minuty - i co ważne, nie pomija notatek które ja bym przeoczył bo zapomniałem o ich istnieniu. ### Q&A - odpowiedzi z cytowaniami w 30 sekund "What do my notes say about context engineering?" - agent czyta indeksy, identyfikuje kandydatów, czyta notatki, syntetyzuje odpowiedź z cytowaniami. Dobre odpowiedzi mogą trafić z powrotem do wiki jako nowe compiled-notes. To zamienia bazę wiedzy z pasywnego archiwum w **aktywne narzędzie myślenia**. Zamiast grzebać w folderach, zadaję pytanie i dostaję syntezę z moich własnych notatek - z linkami do źródeł. Co kluczowe: agent odpowiada na podstawie MOJEJ wiedzy, nie ogólnego internetu. Jeśli trzy miesiące temu zapisałem ważny insight z konferencji - Q&A go znajdzie i przywołuje, nawet jeśli ja sam zapomniałem że go zapisałem. Inny przykład: "Compare what my notes say about RAG vs what Karpathy describes as LLM Wiki." Agent czyta obie notatki, identyfikuje punkty zbieżności i rozbieżności, prezentuje porównanie. To typ analizy, który ręcznie wymagałby otwarcia dwóch dokumentów i porównywania paragraf po paragrafie. ### LINT - health-check Periodyczny health-check vault: broken wikilinks, orphan notes (notatki bez żadnych powiązań), brakujące summaries, TODO markery, niekompletny frontmatter. Agent generuje raport do `_outputs/reports/` i proponuje fixy. Typowy raport LINT po miesiącu: - 4 broken wikilinks (notatki zmienione/przeniesione bez aktualizacji referencji) - 2 orphan notes (dodane ale nigdy nie zlinkowane z resztą vault) - 7 brakujących summaries (notatki z pustym polem summary we frontmatter) - 3 notatki z TODO markerami (niedokończone przetwarzanie) LINT to workflow, o którym nie myślisz aż nie zobaczysz raportu z 23 broken wikilinks po miesiącu dodawania notatek. To hygiene - jak linting kodu. Nie jest sexy, ale bez niego vault degeneruje się w ciszy. Agent robi to za mnie - regularnie i bez zapominania. ## Od ręcznych notatek do 185 plików pilnowanych przez agenta Ewolucja tego systemu nie była planowana. Nie usiadłem w 2022 roku z architekturą w głowie. To był organiczny proces, w którym każda faza rozwiązywała konkretny problem z poprzedniej. **Faza 1 (2022-2024): Ręczne notatki w Obsidian.** Klasyczny setup - foldery tematyczne, wikilinki, tagi. Zapisywałem notatki z książek, konferencji, kursów. Działało do ~50 notatek. Potem zaczął rosnący chaos: niepełne cross-referencje, zapomniany content, duplikaty. Szukanie konkretnej informacji zaczęło trwać dłużej niż po prostu wyszukanie jej w Google od nowa. Typowa trajektoria PKM - entuzjazm, plateau, frustracja. **Faza 2 (2024-2025): Agent zaczyna pomagać.** Kiedy zacząłem intensywnie pracować z Claude Code, naturalnie zacząłem go używać do prostych tasków w vault. Najpierw formatowanie i uzupełnianie frontmatter - nudna robota, idealna dla agenta. Potem linkowanie nowych notatek z istniejącymi - agent jest w tym lepszy ode mnie, bo czyta cały catalog za każdym razem. W końcu - pełne przetwarzanie źródeł od zera: drop plik, powiedz `ingest`, agent robi resztę. W tej fazie powstał pierwszy CLAUDE.md - jeszcze prosty, ~80 linii, ale już definiujący strukturę i podstawowe zasady. **Faza 3 (2026): Karpathy LLM Wiki → formalizacja.** Wątek Karpathy'ego dał mi framework do nazwania tego co miałem. Trzy warstwy - raw sources, wiki, schema - nagle miały oficjalne nazwy. Navigation protocol przestał być "sposób w jaki to robię" i stał się udokumentowaną procedurą. Safety rules, które trzymałem w głowie, trafiły do CLAUDE.md. Sformalizowałem procesy, dopisałem brakujące elementy, uporządkowałem indeksy. CLAUDE.md urosło z 80 do 300+ linii. Stan dziś: **185 notatek**, 13 kategorii tematycznych. Breakdown: - `AI/` - 19 notatek (Claude Code, harness engineering, context engineering, skills) - `LIFE/` - 40 notatek (książki, wiedza, narzędzia) - `BUSINESS/` - 26 notatek - `CODE/` - 23 notatki - `PROJECTS/` - 12 notatek (Qamera AI, Brain, Agentic Systems) Porównanie workflow - kiedyś i teraz: - **Kiedyś:** przeczytam artykuł → może zapiszę link w bookmarkach → zapomnę gdzie i w jakim kontekście → szukam od nowa gdy potrzebuję → tracę 20 minut na odtworzenie kontekstu - **Teraz:** klikam Obsidian Web Clipper → artykuł ląduje w `_raw/inbox/` → piszę `ingest` w Claude Code → agent wplata notatkę w sieć wiedzy z cross-referencjami → przy pytaniu: Q&A w 30 sekund z cytowaniami i linkami do źródeł > _"Wiki to persistent, compounding artifact - cross-references gotowe, contradictions flagged."_ Co się compound-uje: agent pilnuje struktury i standardów, każda notatka automatycznie trafia we właściwe miejsce z właściwymi linkami. Wiedza się kumuluje, nie rozprasza. Im więcej notatek, tym bogatsza sieć powiązań - i tym cenniejsza każda następna notatka, bo agent ma więcej kontekstu do linkowania. To odwrócenie tradycyjnego problemu PKM, gdzie więcej notatek = więcej chaosu. Tutaj więcej notatek = bogatszy graph. ## Notatki ułożone dla agenta, nie dla estetyki To jest kluczowy mindset shift, który zmienił moje podejście do PKM. Nie organizujesz notatek żebyś TY łatwiej szukał. Organizujesz je żeby **AGENT łatwiej znajdował**. Agent jest primary consumer Twojej bazy wiedzy. Ty jesteś curator - odpowiadasz za sourcing (co wrzucasz do inbox), eksplorację (jakie pytania zadajesz) i decyzje (co kompilować, co archiwizować). Agent odpowiada za bookkeeping: tworzenie notatek, aktualizację indeksów, pilnowanie cross-referencji, flagowanie sprzeczności. To podział, który pozwala Ci skupić się na myśleniu, nie na administracji. **Progressive disclosure** minimalizuje zużycie okna kontekstowego. Agent nie musi czytać 185 plików żeby odpowiedzieć na pytanie. Czyta 80 linii vault-map, identyfikuje relevant area, czyta kilka konkretnych notatek. Trafniejsze odpowiedzi, niższy koszt, szybsze działanie. Żeby to zobaczyć w liczbach: gdyby agent czytał cały vault przy każdym zapytaniu, to ~185 plików × ~500 słów = ~92 500 słów wstrzykniętych do context window. Z progressive disclosure: vault-map (80 linii) + 3-5 relevantnych notatek = ~3000 słów. **30x mniej kontekstu, ale trafniejsze odpowiedzi** - bo agent wie co czyta i dlaczego. Paralela z agentic coding jest bezpośrednia: | | Agentic Coding | LLM Knowledge Base | | ----- | ------------------------------------------------- | ------------------------------------ | | Ty | Architekt środowiska (specs, context, guardrails) | Curator (sourcing, pytania, decyzje) | | Agent | Implementuje kod | Bookkeeper, writer, cross-referencer | Agent nie zastępuje myślenia - zastępuje bookkeeping. To nie jest subtelna różnica. Decyzja CO ingestować, jakie pytania zadać, co kompilować - to nadal Twoja robota. Agent jest maszyną do utrzymania porządku, nie maszyną do myślenia za Ciebie. Kiedy próbujesz oddać agentowi decyzje - dostajesz generyczne, "bezpieczne" odpowiedzi. Kiedy oddajesz mu utrzymanie, a sam skupiasz się na kuratorstwie - dostajesz system, który się compound-uje z każdą nową notatką. Każde źródło wzbogaca istniejącą sieć wiedzy. Cross-referencje powstają automatycznie. Sprzeczności między notatkami są flagowane. Summary jest zawsze aktualne. > _"Shift: from 'developer who writes code' to 'architect who designs systems for agents to write code.'"_ Ten sam shift dotyczy wiedzy: from "person who maintains notes" to "curator who designs systems for agents to maintain knowledge." Ty decydujesz co jest warte zapisania. Agent dba o to, żeby zapisane informacje były zawsze dostępne, powiązane i aktualne. ## Co z tego wynika Pięć wniosków po trzech latach budowania tego systemu: 1. **Metodologia > narzędzie.** Zettelkasten, Building a Second Brain, LLM Wiki - to filozofie, nie aplikacje. Obsidian, Notion, Logseq - to narzędzia. Kluczowe jest zrozumienie _dlaczego_ organizujesz wiedzę w określony sposób, nie _w czym_. Możesz zbudować ten sam system w Notion z agentem, albo w czystym VS Code z plikami Markdown. Narzędzie jest wymienne - zasady nawigacji, indeksowania i progressive disclosure działają wszędzie. 2. **CLAUDE.md to kontrakt z agentem.** Ewoluuje z doświadczenia - zaczynasz od prostego, dodajesz na podstawie tego co agent robi źle, usuwasz co nie działa. Po miesiącu masz solidną konfigurację. Po roku - dojrzały system. 3. **Progressive disclosure działa.** Trzy poziomy indeksów (vault-map → catalog → graph) to nie over-engineering. To jedyne co pozwala agentowi działać sensownie na 185+ notatkach bez grep-owania całego katalogu i zaśmiecania context window. 4. **Agent nie zastępuje myślenia.** Zastępuje bookkeeping. To fundamentalna różnica. Kiedy próbujesz oddać agentowi _decyzje_ - dostaniesz generyczne odpowiedzi. Kiedy oddajesz mu _utrzymanie_ - dostaniesz system, który się compound-uje. 5. **Git repo jako fundament.** Version history, branching, collaboration - za darmo. Quartz 4 jako SSG daje Ci publiczny digital garden na GitHub Pages. Cały vault to repozytorium Git - każda zmiana jest tracked, revertable, diffable. Jeśli agent zrobi coś złego (a zrobi, szczególnie na początku) - `git diff` pokaże co zmienił, `git revert` cofnie. To safety net, bez którego nie oddałbym agentowi kontroli nad 185 plikami. Jeśli chcesz zacząć od zera i nie masz jeszcze vault z agentem, przeczytaj [Second Brain z Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills) - tam opisuję jak wystartować od instalacji Obsidian po pierwsze Skills. Ten artykuł pokazuje gdzie możesz dojść po kilku latach iteracji - od prostego vault z kilkunastoma notatkami do systemu z 185 plikami, trzema warstwami i czterema workflow, pilnowanego przez agenta AI. Nie musisz budować tego wszystkiego od razu. Zacznij od CLAUDE.md z 50 liniami, jednego workflow (INGEST) i jednego indeksu (vault-map). Reszta dojdzie organicznie - tak jak doszła u mnie.

Chcesz zbudować podobny system zarządzania wiedzą?

Pomogę Ci zaprojektować architekturę bazy wiedzy z agentem AI - od struktury CLAUDE.md przez indeksy nawigacyjne po workflows dostosowane do Twoich potrzeb.

Umów bezpłatną konsultację
## Zasoby - [Wątek Karpathy'ego na X](https://x.com/karpathy/status/2039805659525644595) - oryginalny post o LLM Wiki (kwiecień 2026) - [LLM Wiki gist na GitHub](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) - konfiguracja i dokumentacja - [brain.lipowczan.pl](https://brain.lipowczan.pl) - mój vault publiczny (Quartz 4 + GitHub Pages) - [Quartz 4](https://quartz.jzhao.xyz/) - Static Site Generator do Obsidian vault - [Obsidian](https://obsidian.md/) - edytor notatek Markdown - Powiązany artykuł: [Second Brain z Obsidian i Claude Code Skills](/blog/second-brain-obsidian-claude-code-skills) - jak zacząć od zera - [second-brain-template](https://github.com/plipowczan/second-brain-template) - gotowy szablon do startu własnego second brain ## FAQ
### Czym różni się LLM Wiki od tradycyjnego RAG na dokumentach? RAG (Retrieval-Augmented Generation) szuka fragmentów tekstu na żądanie - dostajesz odpowiedź opartą o najbliższe embeddings, ale bez szerszego kontekstu. LLM Wiki to persistent baza wiedzy, w której agent aktywnie buduje i utrzymuje notatki: tworzy cross-referencje, syntetyzuje wiedzę z wielu źródeł i flaguje sprzeczności. Nie odpowiada na pojedyncze pytanie - utrzymuje kompletny, ewoluujący system wiedzy.
### Czy potrzebuję Obsidian żeby zbudować podobny system zarządzania wiedzą z AI? Nie - Obsidian to wygodny frontend z graph view i pluginami, ale nie jest wymagany. Wystarczy dowolny edytor tekstowy + Git repo z plikami `.md`. Kluczowe elementy to CLAUDE.md (schema document definiujący jak agent ma traktować vault) i indeksy nawigacyjne (vault-map, catalog, graph). Możesz zbudować ten sam system w VS Code, Cursor, albo nawet w terminalu.
### Ile czasu zajmuje napisanie CLAUDE.md dla knowledge base i jak zacząć? Pierwsza wersja zajmuje 1-2 godziny - definiujesz strukturę katalogów, podstawowy frontmatter schema i jeden workflow (najlepiej INGEST). Ale CLAUDE.md to żywy dokument, który ewoluuje z doświadczenia. Zaczynasz od prostych reguł, obserwujesz co agent robi źle, dodajesz korekty. Po miesiącu regularnego używania masz solidną konfigurację. Kluczowa zasada: pisz ręcznie, nie generuj przez LLM.
### Czym jest progressive disclosure i dlaczego jest kluczowe dla bazy wiedzy z agentem AI? Progressive disclosure to zasada organizacji notatek w warstwach - od ogólnego spisu treści (vault-map, ~80 linii) przez katalog (catalog, ~650 linii) do konkretnych plików. Agent czyta tylko tyle ile potrzebuje, nie zaśmiecając okna kontekstowego niepotrzebnymi informacjami. To kluczowa różnica między bazą wiedzy "która działa" a taką, gdzie agent gubi się w 185 plikach i zwraca generyczne odpowiedzi.
### Czy system LLM Wiki wymaga technicznego background'u do wdrożenia? Podstawowy setup wymaga znajomości Git, terminala i plików Markdown - nie musisz kodować. Najtrudniejsza część to napisanie dobrego CLAUDE.md, które precyzyjnie definiuje jak agent powinien traktować Twój vault. Ten artykuł opisuje dojrzały system po 3 latach iteracji, ale zacząć możesz prosto - od jednego folderu, prostego frontmatter schema i workflow INGEST. Artykuł [Second Brain z Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills) daje solidną bazę startową.
--- # 5 repozytoriów GitHub, które zmienią Twoją pracę z Claude Code Source: https://pawel.lipowczan.pl/blog/5-repozytoriow-github-claude-code Published: 2026-03-31 # 5 repozytoriów GitHub, które zmienią Twoją pracę z Claude Code Używanie Claude Code bez ekosystemu skills to jak korzystanie ze smartfona bez aplikacji. Niby działa, ale zostawiasz na stole większość potencjału. Sam przez to przechodziłem - otwierałem terminal, wpisywałem prompty, dostawałem kod. Czasem dobry, czasem generyczny. Zero systemu, zero powtarzalności. Potem zacząłem eksplorować GitHub i odkryłem, że wokół Claude Code wyrósł potężny ekosystem. Skills, frameworki, integracje - setki repozytoriów, z których każde obiecuje rewolucję. Problem? Większość to szum. Trudno oddzielić narzędzia, które naprawdę robią różnicę, od tych, które wyglądają dobrze w README, ale nie sprawdzają się w codziennej pracy. W tym artykule zebrałem **5 repozytoriów GitHub**, które naprawdę zmieniły sposób, w jaki pracuję z Claude Code. Każde z nich przetestowałem w produkcyjnym workflow - nie polecam rzeczy, których sam nie używam. ## 1. UI/UX Pro Max - koniec z generycznym AI slop Znasz ten problem. Prosisz Claude Code o stworzenie strony i dostajesz... dokładnie to samo co wszyscy inni. Ten sam layout z hero section, te same zaokrąglone karty, te same gradienty. **Generic AI slop** - generyczny wygląd, który natychmiast zdradza, że stronę wygenerował AI. **UI/UX Pro Max** to skill, który rozwiązuje ten problem u źródła. Zamiast jednego uniwersalnego podejścia do designu, oferuje **inteligentną generację design system'ów** dopasowanych do tego, co faktycznie budujesz. ### Jak to działa Skill analizuje typ projektu - portfolio, SaaS, e-commerce, landing page - i dobiera odpowiedni system designu. Inne kolory, inne proporcje, inne komponenty. To nie jest losowy wybór. Każdy system ma swoją logikę: portfolio podkreśla osobistą markę, SaaS kładzie nacisk na konwersję, e-commerce na prezentację produktów. W praktyce oznacza to, że dwa różne projekty wygenerowane z tym samym skillem wyglądają **zupełnie inaczej**. A to dokładnie o to chodzi - indywidualność zamiast szablonu. ### Moje doświadczenie Używam UI/UX Pro Max w połączeniu z **Tailwind CSS** i **React**. Przy budowie komponentów dla klientów skill generuje spójny design system, który potem dostosowuję. Oszczędza mi to czas na etapie prototypowania - zamiast zaczynać od zera lub walczyć z generycznym outputem, mam solidną bazę dopasowaną do kontekstu projektu. Jeśli budujesz cokolwiek z frontendem i chcesz, żeby wyglądało profesjonalnie bez zatrudniania designera - to jest Twój punkt startu. **Repozytorium:** [UI/UX Pro Max](https://github.com/nextlevelbuilder/ui-ux-pro-max-skill) ## 2. OpenSpec - strukturyzowany development zamiast chaosu **OpenSpec (OPSX)** to framework do spec-driven development, który naprawdę zmienił mój sposób pracy z Claude Code. Używam go codziennie. ### Problem, który rozwiązuje Każdy, kto pracuje z Claude Code dłużej niż tydzień, zna ten scenariusz: zaczynasz sesję, budujesz feature, kontekst rośnie, agent zaczyna "zapominać" wcześniejsze ustalenia. To **context window rot** - degradacja jakości odpowiedzi w miarę jak konwersacja się wydłuża. OpenSpec rozwiązuje to przez **spec-driven development**. Zamiast chaotycznych sesji, gdzie mówisz agentowi, co ma robić krok po kroku, tworzysz strukturyzowane artefakty: specyfikację zmiany, plan implementacji, delta specs. Agent wie, co buduje, dlaczego i jak - zanim napisze pierwszą linijkę kodu. ### Jak wygląda workflow Typowa sesja z OpenSpec zaczyna się od **explore** - i to jest kluczowy krok, który większość osób pomija: ```text 0. /opsx:explore → brainstorming z agentem 1. /opsx:new "dodaj dark mode do bloga" → tworzy change z artefaktami 2. /opsx:ff → fast-forward przez wszystkie artefakty 3. /opsx:apply → implementacja zadań z planu 4. /opsx:verify → weryfikacja vs specyfikacja 5. /opsx:archive → archiwizacja ukończonej zmiany ``` Faza explore to moment, w którym agent odbija z Tobą pomysły, dopytuje o szczegóły, proponuje podejścia. To tutaj zapada decyzja - czy w ogóle potrzebujesz pełnej specyfikacji, czy wystarczy sam proposal z listą zadań. Prosta zmiana nie wymaga rozbudowanej specyfikacji. Złożony feature - jak najbardziej. Z mojego doświadczenia: **im więcej czasu spędzisz na etapie przygotowania dobrej specyfikacji, tym mniej iteracji będziesz potrzebować przy samej implementacji kodu**. To się zwraca wielokrotnie. Każdy krok produkuje konkretny artefakt - plik markdown w repozytorium. Nie tracisz kontekstu między sesjami, bo specyfikacja jest w plikach, nie w historii czatu. ### Co wyróżnia OpenSpec OpenSpec wyróżnia się na tle innych frameworków z kilku powodów: - **Artefakty w repo** - wszystko pod version control, mogę wrócić do specyfikacji po tygodniu - **Delta specs** - zmiany opisane inkrementalnie, łatwo śledzić co się zmieniło - **Integracja z walidacją** - po implementacji mogę zweryfikować, czy kod zgadza się ze specyfikacją Napisałem o tym szczegółowo w osobnym artykule: [OpenSpec - strukturyzowana praca z AI](/blog/opsx-workflow-strukturyzowana-praca-z-ai). **Repozytorium:** [OpenSpec](https://github.com/Fission-AI/OpenSpec/) | **Strona:** [openspec.dev](https://openspec.dev/) ## 3. Excalidraw - diagramy i mapowanie procesów z AI Komunikacja wizualna to jeden z najbardziej niedocenianych aspektów pracy z AI. Możesz opisać architekturę systemu w tysiącu słów - albo narysować jeden diagram. ### Dlaczego Excalidraw Próbowałem różnych podejść. Zaczynałem od **Mermaid** - wyglądało średnio, tekstowa składnia nie pozwalała oddać złożoności procesu. Potem testowałem **Miro** i generowanie diagramów z poziomu języka naturalnego. Problem? Nie dało się nadać outputowi zdefiniowanych styli, charakteru ani dodatkowych informacji kontekstowych. Wyniki były generyczne i wymagały tyle ręcznej pracy, że tracił się sens automatyzacji. Szukałem dalej, aż trafiłem na **Excalidraw** - narzędzie do tworzenia diagramów w stylu hand-drawn: procesów, architektur, flowchartów. Dzięki integracjom z Claude Code możesz generować je bezpośrednio z terminala. ### Integracja z narzędziami, których faktycznie używam Kluczowa przewaga Excalidraw to integracja z ekosystemem, w którym na co dzień pracuję. Plugin do **Obsidian** pozwala przeglądać i edytować diagramy bezpośrednio w vault'cie - tam, gdzie trzymam całą bazę wiedzy. Extension do **Visual Studio Code** daje to samo w IDE, gdzie spędzam większość czasu z agentami AI. To ważne, bo sam Claude Code generuje diagramy, ale ich nie wyświetla. Potrzebujesz narzędzia, które pozwoli Ci nie tylko wygenerować plik `.excalidraw`, ale też go obejrzeć, zmodyfikować i osadzić w kontekście projektu. Obsidian i VS Code to umożliwiają. ### Dwa zastosowania, dwa skills Korzystam z Excalidraw na dwa sposoby: **1. Tłumaczenie koncepcji technicznych** Skill od Cole'a Medina ([excalidraw-diagram-skill](https://github.com/coleam00/excalidraw-diagram-skill)) pozwala prosić Claude Code o wizualizację koncepcji. "Narysuj architekturę tego systemu" albo "pokaż flow danych w tym pipeline" - i dostajesz czytelny diagram zamiast ściany tekstu. **2. Mapowanie procesów dla klientów** Do tego używam zestawu skills z repozytorium [shared-skills](https://github.com/200iqlabs/shared-skills). Kiedyś mapowanie procesu dla klienta oznaczało godziny ręcznej pracy w Miro - rysowanie każdego kroku, łączenie strzałkami, formatowanie. Teraz Claude Code generuje mapę procesu automatycznie na podstawie opisu. Wymaga czasem korekty, ale odchodzi ogrom manualnej roboty. Klient dostaje wizualną dokumentację, a nie listę kroków w markdownie. Więcej o mojej metodologii mapowania procesów w kontekście szukania optymalizacji pisałem w artykule [Każda firma działa nieoptymalnie](/blog/kazda-firma-dziala-nieoptymalnie). ### Wartość w praktyce Diagram wart jest tysiąca słów - dosłownie. Kiedy tłumaczysz klientowi architekturę systemu lub omawiasz z zespołem flow nowej feature'ki, jeden dobry diagram zastępuje godzinę wyjaśnień. A fakt, że mogę go wygenerować bez opuszczania terminala, to game changer. **Excalidraw:** [excalidraw.com](https://plus.excalidraw.com/) | **Diagram Skill:** [GitHub](https://github.com/coleam00/excalidraw-diagram-skill) | **Shared Skills:** [GitHub](https://github.com/200iqlabs/shared-skills) ## 4. Obsidian Skills - pamięć długoterminowa dla agenta AI Claude Code ma pewien fundamentalny problem: **nie pamięta**. Każda nowa sesja zaczyna się od zera. Owszem, masz `CLAUDE.md` i pliki w `.claude/`, ale to nie to samo, co prawdziwa baza wiedzy, do której agent może sięgnąć w dowolnym momencie. **Obsidian Skills** to zestaw narzędzi łączących Claude Code z **Obsidian** - jednym z najlepszych edytorów notatek opartych na plikach markdown. Połączenie tych dwóch narzędzi tworzy coś, co można nazwać **second brain dla AI**. ### Jak to działa Obsidian przechowuje notatki jako zwykłe pliki `.md` na Twoim komputerze. Claude Code ma bezpośredni dostęp do systemu plików. Połącz jedno z drugim i nagle Twój agent ma dostęp do: - **Notatek z projektów** - kontekst, decyzje, lessons learned - **Bazy wiedzy** - dokumentacja, procesy, procedury - **Szablonów** - powtarzalne struktury dokumentów - **Historii** - co robiłeś wczoraj, tydzień temu, miesiąc temu To nie jest MCP server, który ładuje wszystko do context window z góry. Skills ładują się dynamicznie - **progressive disclosure** oznacza, że agent sięga po informację dopiero gdy jej potrzebuje. ### Moje doświadczenie Używam Obsidian jako centrum zarządzania wiedzą. Notatki ze spotkań, plany projektów, research - wszystko trafia do vault'a. Claude Code przetwarza te notatki, tworzy podsumowania, łączy informacje z różnych źródeł. Szczegółowo opisałem ten setup w artykule [Second Brain z Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills). Jeśli szukasz sposobu na to, żeby Twój agent AI faktycznie "wiedział" więcej niż to, co jest w bieżącej sesji - zacznij od tego. **Repozytorium:** [Obsidian Skills](https://github.com/kepano/obsidian-skills) ## 5. Awesome Claude Code - one-stop shop na start Nie wiesz od czego zacząć? **Awesome Claude Code** to odpowiedź. To starannie wyselekcjonowana lista najlepszych zasobów dla Claude Code - skills, workflows, MCP servers, prompts, narzędzia. Jeden punkt wejścia zamiast przeszukiwania setek repozytoriów. ### Co znajdziesz Repozytorium jest podzielone na kategorie: - **Skills** - gotowe skills do instalacji (od designu po testowanie) - **MCP Servers** - integracje z zewnętrznymi usługami - **Workflows** - sprawdzone procesy pracy z Claude Code - **Prompts** - szablony promptów na różne okazje - **Community** - linki do społeczności, tutoriali, artykułów ### Dlaczego to ważne Ekosystem Claude Code rośnie szybko. Nowe skills i narzędzia pojawiają się codziennie. **Awesome Claude Code** oszczędza Ci czas na research - ktoś już przefiltrował dostepne zasoby i zebrał najlepsze w jednym miejscu. To idealne repozytorium na start. Przejrzyj listę, znajdź 2-3 rzeczy, które pasują do Twoich potrzeb, zainstaluj i przetestuj. Potem wróć po więcej. Wiele narzędzi z mojej listy - UI/UX Pro Max, Obsidian Skills - możesz znaleźć właśnie przez Awesome Claude Code. To jak indeks do całego ekosystemu. **Repozytorium:** [Awesome Claude Code](https://github.com/hesreallyhim/awesome-claude-code) ## Jak wybrać - mapa decyzyjna Nie instaluj wszystkich pięciu naraz. Zacznij od jednego, przetestuj, dodaj kolejne, gdy poczujesz potrzebę. | Twoja sytuacja | Repozytorium | Dlaczego | |---|---|---| | Budujesz frontend i chcesz lepszy design | **UI/UX Pro Max** | Koniec z generycznym AI slop | | Zaczynasz nowy projekt lub feature | **OpenSpec** | Struktura zamiast chaosu | | Potrzebujesz diagramów i wizualizacji | **Excalidraw** | Komunikacja wizualna z terminala | | Chcesz pamięć między sesjami | **Obsidian Skills** | Second brain dla AI | | Nie wiesz od czego zacząć | **Awesome Claude Code** | Wyselekcjonowana lista na start | Jeśli musiałbym wybrać tylko jedno - zacząłbym od **OpenSpec**. Strukturyzowane podejście do development zmienia wszystko. Design, diagramy i pamięć to nadbudówki - ale bez dobrego fundamentu workflow, żadne narzędzie nie pomoże. ## Kluczowe wnioski 1. **Claude Code "vanilla" to dopiero początek** - ekosystem skills zmienia zasady gry, dosłownie mnożąc możliwości agenta 2. **Nie instaluj wszystkiego naraz** - wybierz 1-2 repo pasujące do Twojego aktualnego problemu i przetestuj w praktyce 3. **Skills > MCP servers dla większości use cases** - progressive disclosure zapobiega context bloat, ładujesz tylko to co potrzebujesz 4. **Testuj osobiście** - każdy workflow jest inny, moja lista ≠ Twoja lista 5. **Community jest kluczowe** - najlepsze narzędzia powstają w open source, obserwuj repozytoria i bądź na bieżąco ---

Chcesz skonfigurować Claude Code pod swój workflow?

Pomogę Ci wybrać odpowiednie skills i narzędzia, skonfigurować środowisko agentowe i zbudować workflow, który naprawdę przyspieszy Twoją pracę.

Umów bezpłatną konsultację
## Przydatne zasoby - [UI/UX Pro Max](https://github.com/nextlevelbuilder/ui-ux-pro-max-skill) - skill do inteligentnej generacji design system'ów - [OpenSpec](https://github.com/Fission-AI/OpenSpec/) - spec-driven development framework ([strona](https://openspec.dev/)) - [Excalidraw Diagram Skill](https://github.com/coleam00/excalidraw-diagram-skill) - generowanie diagramów z Claude Code - [Shared Skills (200iqlabs)](https://github.com/200iqlabs/shared-skills) - skills do mapowania procesów - [Obsidian Skills](https://github.com/kepano/obsidian-skills) - integracja Claude Code z Obsidian - [Awesome Claude Code](https://github.com/hesreallyhim/awesome-claude-code) - wyselekcjonowana lista zasobów - [Claude Code Documentation](https://docs.anthropic.com/claude-code) - oficjalna dokumentacja - [5 technik pracy z Claude Code](/blog/5-technik-pracy-z-claude-code) - powiązany artykuł - [OpenSpec - strukturyzowana praca z AI](/blog/opsx-workflow-strukturyzowana-praca-z-ai) - pełny artykuł o OPSX - [Second Brain z Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills) - artykuł o second brain - [Środowisko agentowe AI](/blog/srodowisko-agentowe-ai-dwie-firmy) - architektura multi-agent ## FAQ
### Czy te repozytoria działają z najnowszą wersją Claude Code i są aktywnie rozwijane? Tak, wszystkie pięć repozytoriów jest aktywnie rozwijanych i kompatybilnych z aktualną wersją Claude Code. Przed instalacją warto sprawdzić datę ostatniego commitu na GitHub - ekosystem zmienia się szybko i pojawiają się nowe wersje. Skills instalujesz jako pliki w repozytorium, więc nie ma ryzyka złamania kompatybilności jak przy aktualizacji zależności.
### Czy mogę używać kilku skills jednocześnie w jednym projekcie bez problemów z wydajnością? Tak, skills działają na zasadzie progressive disclosure - ładują się dynamicznie, tylko gdy są potrzebne. Możesz mieć zainstalowane dziesiątki skills, a context window nie będzie zaśmiecony. To kluczowa różnica w porównaniu z MCP servers, które ładują wszystkie narzędzia z góry. W praktyce w jednym projekcie używam równocześnie OpenSpec, UI/UX Pro Max i kilku innych bez żadnych problemów.
### Czy UI/UX Pro Max zastępuje wiedzę o designie i doświadczenie w CSS? Nie zastępuje, ale wyrównuje szanse. Programista bez doświadczenia w UI dostanie spójny, profesjonalny design dopasowany do typu projektu zamiast generycznego szablonu. Jeśli masz doświadczenie w designie, skill przyspieszy Twoją pracę - dostajesz solidną bazę do dalszej customizacji. Traktuj go jak inteligentny starter kit, nie jak zamiennik designera.
### Czym różni się OpenSpec od innych frameworków do pracy z AI? OpenSpec to spec-driven development - najpierw tworzysz strukturyzowaną specyfikację zmiany, potem implementujesz. Artefakty (pliki markdown) żyją w repozytorium pod version control, więc nie tracisz kontekstu między sesjami. OpenSpec wyróżnia się fazą explore (brainstorming przed implementacją), delta specs (inkrementalne opisy zmian) i wbudowaną weryfikacją implementacji vs specyfikacja.
### Czy Excalidraw wymaga płatnej subskrypcji, żeby działać z Claude Code? Nie, Excalidraw jest open source i darmowy. Excalidraw+ oferuje dodatkowe funkcje jak real-time collaboration, ale skills do Claude Code działają z darmową wersją. Generujesz diagramy lokalnie jako pliki, bez potrzeby konta czy subskrypcji. Jedyne co potrzebujesz to zainstalowany skill i Claude Code.
### Od którego repozytorium powinienem zacząć, jeśli dopiero zaczynam pracę z Claude Code? Zacznij od Awesome Claude Code - to wyselekcjonowana lista, z której możesz wybrać narzędzia pasujące do Twoich konkretnych potrzeb. Potem dodaj OpenSpec do strukturyzowania pracy - to fundament, który zmienia sposób interakcji z agentem. Resztę dodawaj stopniowo, gdy pojawi się realna potrzeba. Nie instaluj wszystkiego na start - lepiej opanować jedno narzędzie dobrze niż pięć powierzchownie.
--- # Moje środowisko agentowe - jak buduję AI OS dla dwóch firm Source: https://pawel.lipowczan.pl/blog/srodowisko-agentowe-ai-dwie-firmy Published: 2026-03-23 # Moje środowisko agentowe - jak buduję AI OS dla dwóch firm Siadam rano do komputera, otwieram terminal i mówię: "sprawdź stan konta firmowego i porównaj z planem finansowym". Agent CFO ładuje kontekst, łączy się z Revolut API, analizuje ostatnie faktury i zostawia trzy rekomendacje. Potem pytam innego agenta o monitorowanie konkurencji - i dostaję briefing. Trzeci sprawdza terminy prawne i przypomina o zbliżającym się deadline umowy z klientem. Każdy z tych agentów działa pod moim nadzorem - świadomie nie puszczam ich jeszcze w pełni autonomicznie. Uważam, że na tym etapie warto kilka razy przejść operację pod kontrolą, zanim ustawisz scheduler do działań cyklicznych. Ale sam fakt, że mam **osiem wyspecjalizowanych agentów**, z których każdy zarządza innym obszarem - finansami, prawem, marketingiem, contentem, produktem - to już ogromna zmiana. I co najważniejsze - ten sam zestaw agentów obsługuje **dwie różne firmy jednocześnie**. W artykule o [Skills 2.0](/blog/skills-2-0-multi-agent-system-zarzadzanie-firma) opisywałem jak zbudowałem system wieloagentowy. Teraz pokażę coś głębszego - **architekturę**, która za tym stoi. Bo to nie agenty są rewolucyjne. Rewolucyjna jest separacja warstw, która sprawia, że cały system jest przenośny, wersjonowany i niezależny od dostawcy AI. W tym artykule zobaczysz: - Architekturę trzech warstw (Skills → Context → Tools) i dlaczego ta separacja jest kluczowa - Jak ten sam zestaw agentów obsługuje dwie firmy z kompletnie różnymi kontekstami - Dlaczego Git jest fundamentem zaufania do AI - i dlaczego bez niego nie dałbym agentom swobody - Jak moje podejście wypada na tle alternatyw: Perplexity Computer, OpenClaw, Claude Dispatch ## Dwie firmy, jeden system Pracuję na co dzień z dwoma repozytoriami-hubami, które nie są typowymi projektami kodu. To **systemy doradcze oparte na agentach AI**, zbudowane wokół plików Markdown i skryptów CLI. ### 200IQ Labs - spółka technologiczna **200IQ Labs** to prosta spółka akcyjna, która rozwija produkt **Qamera AI** (wirtualne studio fotograficzne AI dla e-commerce) i oferuje wdrożenia środowisk agentowych w firmach. Repozytorium `agentic-ai-system` zawiera pełny kontekst firmowy - finanse, zespół, marka, operacje, klienci. Wszystko w Markdown z nagłówkami "Last updated", żeby agenty wiedziały czy dane są aktualne. Agenci w tym repo to: **CFO** (z integracjami Revolut Business API, Stripe, inFakt), **Doradca Podatkowy**, **Prawnik**, **Konsultant Biznesowy**, **Product Manager**, **LinkedIn Content** i **Marketing**. Każdy ma swój `SKILL.md` z instrukcjami, referencjami i narzędziami. ### PLSoft - jednoosobowa działalność **PLSoft** to moja JDG aktywna od 2008 roku - szkolenia, doradztwo i wdrożenia z zakresu automatyzacji i AI. Repozytorium `agentic-ai-private` ma tę samą architekturę, ale z innym kontekstem biznesowym. Te same **shared-skills** (CFO, Tax Advisor, Legal, Business Consultant) działają tu z danymi PLSoft zamiast 200IQ Labs. Do tego dochodzi unikalny skill **Coach The Five** - coaching biznesowy oparty na metodologii Tomasza Karwatki dla pierwszych 5 lat firmy. Jest też kontekst dla newslettera **Tech News Weekly** (~700 subskrybentów) i marki osobistej na LinkedIn. **Kluczowy insight:** te same agenty, różne konteksty. Agent CFO wie jak analizować finanse - to jest skill. Ale **jakie** finanse analizuje - to jest context. Ta separacja jest fundamentem całej architektury. ## Architektura trzech warstw Oba repozytoria opierają się na tym samym wzorcu architektonicznym. Wyobraź sobie to jako stos: ```text ┌─────────────────────────────────────┐ │ IDE (Claude Code / Cursor / etc.) │ ← Interfejs użytkownika ├─────────────────────────────────────┤ │ Skills (SKILL.md + references/) │ ← Wiedza domenowa (przenośna) ├─────────────────────────────────────┤ │ Context (context/*.md) │ ← Dane firmowe (unikalne per firma) ├─────────────────────────────────────┤ │ Tools (scripts CLI) │ ← Integracje z API └─────────────────────────────────────┘ ``` Każda warstwa ma jasno określoną odpowiedzialność i jest niezależna od pozostałych. Mogę podmienić IDE bez zmiany skills. Mogę dać klientowi te same skills z jego danymi firmowymi. Mogę dodać nowe narzędzie bez modyfikacji wiedzy domenowej. ### Skills - wiedza domenowa **Skill** to modułowa instrukcja w pliku `SKILL.md` z katalogiem `references/` zawierającym szczegółowe materiały referencyjne. Struktura wygląda tak: ```yaml --- name: cfo description: >- Financial advisor and fractional CFO for business analysis. Use when analyzing cash flow, runway, costs, profitability. --- # CFO Agent ## Quick Reference | Aspekt | Wartość | |------------|--------------------------------| | Rola | Dyrektor finansowy | | Integracje | Revolut, Stripe, inFakt | | Triggery | finanse, koszty, przychody | ## Workflow 1. Sprawdź aktualność danych (Last updated) 2. Załaduj kontekst finansowy 3. Wykonaj analizę 4. Przygotuj rekomendacje ``` Kluczowe cechy skills: - **Przenośność** - skill CFO działa identycznie w 200IQ Labs i PLSoft. Zmienia się tylko kontekst (dane firmowe), nie wiedza (jak analizować finanse). - **Progressive disclosure** - `references/` ładowane są dopiero gdy rozmowa tego wymaga. Oszczędza context window, który jest ograniczony i drogi. - **Podział na shared i private** - `shared-skills` (Apache 2.0, open-source) to wiedza domenowa, którą mogę udostępnić klientom. `private-skills` to proprietary logic specyficzny dla mojego biznesu. ### Context - dane firmowe Kontekst to dane **unikalne dla każdej firmy**. Struktura katalogów wygląda tak: ```text context/ ├── company/ │ ├── overview.md # Misja, struktura, dane rejestrowe │ ├── team.md # Zespół, role, kompetencje │ └── operations.md # Procesy operacyjne ├── finance/ │ ├── current-state.md # Aktualne salda, runway │ ├── budget-2026.md # Plan finansowy │ └── revenue-streams.md # Źródła przychodów ├── clients/ │ ├── active/ # Aktywni klienci │ └── pipeline/ # Potencjalni klienci └── brand/ ├── tone-of-voice.md # Styl komunikacji └── linkedin-strategy.md # Strategia social media ``` Każdy plik Markdown zawiera nagłówek **"Last updated: YYYY-MM-DD"**. To minimum - agent widzi czy dane mają tydzień czy trzy miesiące i może ostrzec, że kontekst wymaga aktualizacji. ### Tools - integracje API Trzecia warstwa to **lekkie skrypty CLI** (bash/Python) pobierające dane na żywo z zewnętrznych systemów. Świadomie wybrałem to podejście zamiast ciężkich frameworków MCP. ```bash #!/bin/bash # tools/revolut-balance.sh - pobierz aktualne salda z Revolut Business API curl -s -H "Authorization: Bearer $REVOLUT_TOKEN" \ "https://b2b.revolut.com/api/1.0/accounts" \ | jq '.[] | {currency, balance}' ``` Dlaczego lekkie skrypty zamiast MCP? Trzy powody: 1. **Mniejsze zużycie context window** - definicje narzędzi MCP potrafią zajmować 40 000+ tokenów. Skrypt CLI to kilka linii. 2. **Łatwiejsze debugowanie** - `bash -x tools/revolut-balance.sh` i widzisz dokładnie co się dzieje. 3. **Zero zależności** - bash i curl są wszędzie. Nie potrzebuję specjalnych frameworków. To nie znaczy, że MCP jest zły - dla dużych organizacji z dziesiątkami integracji ma sens. Ale dla mojej skali lekkie skrypty wygrywają pragmatyzmem. ## Jeden system, wiele IDE Jednym z celów projektowych było **uniezależnienie od konkretnego IDE**. Skills w formacie Markdown działają w dowolnym narzędziu, które czyta pliki. Ale żeby to działało w praktyce, potrzebna jest synchronizacja. Architektura opiera się na **Git submodules** i **symlinkach**: - `shared-skills/` - submoduł Git z open-source skills (CFO, Legal, Tax, Business Consultant) - `private-skills/` - osobne repozytorium z proprietary skills - Symlinki do katalogów IDE: `.claude/skills/`, `.github/copilot/`, `.cursor/skills/`, `.agent/skills/` Automatyczna synchronizacja odbywa się przez skrypt i git hooks: ```bash #!/bin/bash # tools/sync-skills.sh - synchronizuj skills do wszystkich IDE SKILLS_DIR="shared-skills" TARGETS=(".claude/skills" ".github/copilot" ".cursor/skills") for target in "${TARGETS[@]}"; do mkdir -p "$target" for skill in "$SKILLS_DIR"/*/; do skill_name=$(basename "$skill") ln -sfn "../../$skill" "$target/$skill_name" done done echo "Skills synced to ${#TARGETS[@]} IDE targets" ``` Git hooks (`post-checkout`, `post-merge`) automatycznie uruchamiają ten skrypt po aktualizacji submodułów. Efekt? Aktualizujesz skill w jednym miejscu - zmiana propaguje się do Claude Code, Cursor, Copilot i Antigravity jednocześnie. ### Auto-triggering Nie musisz ręcznie mówić "teraz chcę rozmawiać z CFO". Agenty **aktywują się automatycznie** na podstawie słów kluczowych w zapytaniu. Napisz "ile mamy na koncie?" albo "przeanalizuj koszty" - a agent CFO włącza się sam, ładuje kontekst finansowy i odpowiada z pozycji dyrektora finansowego. To działa dzięki polu `description` w metadanych skilla. Dobre opisy z precyzyjnymi triggerami ("cash flow", "runway", "koszty", "przychody") sprawiają, że orkiestrator wybiera właściwego agenta bez interwencji użytkownika. ## Git jako fundament zaufania To jest sekcja, którą uważam za najważniejszą w całym artykule. Bo technologia agentów zmienia się co tydzień. Ale **problem zaufania** zostanie z nami na lata. Ludzie boją się dawać AI dostęp do swoich danych i plików. I mają rację - opisywałem realne zagrożenia w artykule o [OpenClaw](/blog/openclaw-bezpieczenstwo-agentow-ai), gdzie 28 tysięcy instancji było odsłoniętych na internet. Ale rozwiązaniem nie jest ograniczanie AI do bezużyteczności. Rozwiązaniem jest **budowanie systemów kontroli**. Git daje mi dokładnie to: - **`git diff`** - widzę dokładnie co agent zmienił w każdym pliku - **`git revert`** - cofam w każdej chwili do dowolnego punktu - **`git log`** - pełna historia zmian z timestampami i opisami - **`git blame`** - wiem kto (lub co) zmodyfikowało konkretną linię To tworzy pętlę zaufania: **im większa kontrola → im większa swoboda dla agenta → im szybciej dostarcza wartość → im więcej zyskuję**. Pisałem o podobnym podejściu w kontekście zarządzania wiedzą w artykule o [Second Brain z Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills). Tam chodziło o organizację notatek. Tutaj stawka jest wyższa - chodzi o dane finansowe, prawne i operacyjne dwóch firm. ### Claude.ai Dispatch - świadome ograniczanie uprawnień Oprócz Claude Code używam też funkcji **Dispatch** w Claude.ai, która umożliwia pracę na plikach i podłączonych narzędziach. Ale **świadomie ograniczam autonomię i dostępy**: - **Dostęp do odczytu:** ClickUp (zadania), Revolut (salda), Stripe (subskrypcje), repozytoria GitHub - **Brak dostępu do zapisu** - agent może analizować dane, ale nie może nic modyfikować w tych systemach - **Kontrola zakresu** - precyzyjnie dobieram jakie narzędzia są podłączone Anthropic jest gwarantem bezpieczeństwa środowiska, ale ostateczna decyzja o zakresie dostępu leży po mojej stronie. To fundamentalna różnica wobec podejścia [OpenClaw](/blog/openclaw-bezpieczenstwo-agentow-ai), gdzie domyślnie agent ma pełny dostęp do systemu - i to od użytkownika wymaga się sandboxowania. ## Jak to się ma do rynku Rynek "AI agents" eksplodował w pierwszym kwartale 2026. Żeby osadzić moje podejście w kontekście, porównajmy cztery różne modele: | Aspekt | Moje środowisko | Claude Dispatch | Perplexity Computer | OpenClaw | |--------|-----------------|-----------------|---------------------|----------| | **Model wdrożenia** | Self-managed repos + IDE | Cloud + local hybrid | Fully managed SaaS | Self-hosted runtime | | **Kontrola danych** | 100% lokalna (Git) | Hybrid (Anthropic servers + local) | Cloud provider | 100% lokalna | | **Koszt** | $0 infrastruktury + API calls | $20-200/mies. | $200+/mies. | $0 + API calls | | **Vendor lock-in** | Zero (Markdown, YAML) | Anthropic ecosystem | Perplexity ecosystem | Niski (open source) | | **Bezpieczeństwo** | Git + świadome uprawnienia | Sandbox VM, ~50% niezawodności | Izolowany pod K8s | Wymaga własnego sandboxingu | | **Multi-IDE** | Tak (symlinki) | Nie (tylko Claude) | Nie (web UI) | Nie (własny runtime) | | **Dla kogo** | Tech leaders, developerzy | Użytkownicy Claude | Biznes bez DevOps | Power-userzy, self-hosted | ### Dlaczego wybrałem swoje podejście **Pełna kontrola.** Moje dane nigdy nie opuszczają moich repozytoriów (poza wywołaniami API do modeli). Konteksty firmowe - finanse, klienci, strategia - żyją w Git, nie na serwerach dostawcy. **Przenośność.** Gdyby jutro Claude Code przestał istnieć, moje skills działałyby dalej w Cursor, Copilot lub dowolnym innym narzędziu czytającym Markdown. Zero migracji. **Markdown jako lingua franca.** Uniwersalny format, czytelny zarówno dla ludzi jak i LLM-ów. Bez vendor lock-in. Bez proprietary formats. **Git jako warstwa audytu.** Każda zmiana jest śledzona. Mogę porównać stan przed i po działaniu agenta. Mogę wrócić do dowolnego punktu w czasie. Perplexity Computer jest świetny dla firm, które chcą "plug and play" za $200/mies. OpenClaw daje niesamowitą elastyczność, ale wymaga solidnego DevOps. Claude Dispatch to ciekawy kierunek, ale wciąż research preview. Moje podejście wymaga więcej pracy na starcie, ale daje **pełną kontrolę i zero zależności**. ## Co działa, co nie Po kilku tygodniach intensywnego użytkowania tego systemu na dwóch firmach, mam jasny obraz co się sprawdza, a co wymaga pracy. ### Co działa dobrze 1. **Separacja skills od kontekstu** - mogę dać klientowi ten sam zestaw agentów, on podłącza swoje dane firmowe i działa. Przetestowane na dwóch firmach. 2. **Git jako warstwa bezpieczeństwa** - pełny audyt, rollback, porównywanie zmian. Fundamentalne dla zaufania. 3. **Markdown jako lingua franca** - uniwersalny, czytelny, bez vendor lock-in. Działa w każdym IDE i z każdym modelem. 4. **Lekkie integracje bash/Python** - mniejsze zużycie context window niż MCP, łatwiejsze debugowanie. 5. **Progressive disclosure** - skill ładuje referencje dopiero gdy są potrzebne. Krytyczne przy ograniczonym context window. 6. **Auto-triggering** - agenty aktywują się automatycznie. Nie trzeba ręcznie wybierać "teraz chcę rozmawiać z CFO". ### Co wymaga poprawy 1. **Przenośność pamięci między platformami** - skills działają w wielu IDE, ale pamięć i stan rozmów są locked per platforma. Claude Code nie wie co powiedziałem w Cursor. 2. **Freshness kontekstu** - nagłówki "Last updated" to absolutne minimum. Brakuje automatycznego ostrzegania gdy dane są stare i mechanizmu auto-aktualizacji. 3. **Onboarding klientów** - wdrożenie nowego klienta to ~4h pracy. Za dużo. Skill `environment-setup` pomaga, ale potrzebuję bardziej zautomatyzowanego procesu. 4. **Brak CI/CD dla skills** - weryfikacja jakości skills jest manualna (przez skill-creator evals). Brakuje automatycznego pipeline'u testującego czy skill nadal działa po zmianie. 5. **Sandbox** - Nanoclaw (sandbox od NVIDIA) ma potencjał na autonomiczne zadania nocne (analizy, raporty), ale model bezpieczeństwa wymaga dopracowania zanim dam mu prawdziwe dane. ## Sześć zasad projektowania środowiska agentowego Na podstawie tych doświadczeń wykrystalizowało się sześć zasad, które traktuję jak fundamenty: 1. **Kontrola wersji jest fundamentem** - bez Git nie ma zaufania. Bez zaufania nie ma autonomii dla agenta. Bez autonomii agent jest bezużyteczny. 2. **Separuj wiedzę od danych** - skills (przenośne, open-source) vs context (unikalne per firma, prywatne). To jest ta sama zasada co separation of concerns w programowaniu. 3. **Ograniczaj uprawnienia świadomie** - odczyt tak, zapis z kontrolą, autonomia proporcjonalna do poziomu audytu. Nie dawaj agentowi pełnego dostępu "bo tak wygodniej". 4. **Buduj na otwartych formatach** - Markdown, YAML, skrypty CLI. Zero vendor lock-in. Jutro możesz zmienić dostawcę AI bez zmiany architektury. 5. **Progressive disclosure** - nie ładuj wszystkiego na raz. Context window jest ograniczony i drogi. Skill powinien ładować referencje dopiero gdy są potrzebne. 6. **Code-first, no-code gdy trzeba** - agenty mają być narzędziem dla ludzi technicznych, nie zastępstwem dla nich. Dla klientów bez zespołu technicznego - no-code alternatywy. ## Co dalej To nie jest gotowy produkt - to **żywy system**, który ewoluuje każdego dnia. Kilka kierunków, nad którymi aktywnie pracuję: - **Standaryzacja pamięci agenta** - jak zrobić, żeby pamięć i kontekst rozmów były przenośne między platformami? Dziś Claude Code, Cursor i Copilot mają osobne pamięci. To trzeba rozwiązać. - **Sandbox dla autonomicznych zadań** - Nanoclaw od NVIDIA ma potencjał na nocne analizy i raporty. Ale model bezpieczeństwa wymaga więcej pracy zanim zaufam mu z prawdziwymi danymi. - **Skalowanie na zespół** - wspólne skills, różne konteksty, różne poziomy uprawnień. Jak dać juniorowi read-only dostęp do agenta prawnego, a seniorowi pełne uprawnienia? - **Mierzenie ROI** - ile czasu oszczędzam? Jaka jest jakość decyzji? O ile mniej context-switchingu? Potrzebuję metryk, nie przeczuć. Będę się dzielił postępami. Jeśli budujesz coś podobnego albo rozważasz wdrożenie środowiska agentowego w swojej firmie - opisuję też szczegóły techniczne w artykule o [OPSX Workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai) i [5 technikach pracy z Claude Code](/blog/5-technik-pracy-z-claude-code).

Chcesz zbudować środowisko agentowe dla swojej firmy?

Pomogę Ci zaprojektować architekturę agentów AI dopasowaną do Twojego biznesu - od analizy procesów przez budowę skills po wdrożenie integracji z narzędziami, których już używasz.

Umów bezpłatną konsultację
## FAQ
### Czym jest środowisko agentowe AI i czym różni się od pojedynczego chatbota? Środowisko agentowe to system wielu wyspecjalizowanych agentów AI z dostępem do narzędzi, danych firmowych i integracji z zewnętrznymi systemami. W przeciwieństwie od chatbota, agenty mają persistent memory (pamiętają kontekst między sesjami), automatyczne triggery (aktywują się na słowa kluczowe) i mogą zarządzać konkretnymi obszarami firmy - finansami, prawem, marketingiem. Chatbot odpowiada na pytania, środowisko agentowe **zarządza procesami**.
### Ile kosztuje zbudowanie własnego środowiska agentowego opartego na Agent Skills? Koszty infrastruktury wynoszą $0 - skills i konteksty to pliki Markdown w repozytorium Git, nie wymagają specjalnego hardware'u. Jedyny koszt to API calls do modeli AI (Claude, GPT, Gemini). Przy typowym użyciu biznesowym to $20-200/mies. za API, zależnie od intensywności i wybranego modelu. Do tego dochodzi czas na konfigurację - pierwsze wdrożenie wymaga ~4h pracy technicznej.
### Czy potrzebuję umiejętności programowania, żeby wdrożyć system agentowy w firmie? Tak, podstawowe umiejętności techniczne są potrzebne - obsługa terminala, Git i edycja plików Markdown. To podejście code-first, gdzie agenty są narzędziem dla ludzi technicznych. Dla firm bez zespołu technicznego lepszym wyborem mogą być gotowe rozwiązania SaaS jak Perplexity Computer ($200/mies.) lub Claude Dispatch - dają mniej kontroli, ale zero konfiguracji.
### Jak zapewnić bezpieczeństwo danych firmowych przy pracy z agentami AI? Trzy kluczowe elementy: Git jako warstwa audytu (widzisz dokładnie każdą zmianę agenta w plikach), świadome ograniczanie uprawnień (read-only dostęp do API, zero zapisu w zewnętrznych systemach bez zatwierdzenia) oraz separacja kontekstów (dane firmowe oddzielone od wiedzy domenowej skills). Kontrola wersji eliminuje strach - zawsze możesz cofnąć zmiany przez `git revert`.
### Czy ten system działa tylko z Claude Code, czy mogę użyć Cursor lub innego IDE? System jest celowo niezależny od konkretnego IDE. Skills w formacie Markdown działają w Claude Code, Cursor, GitHub Copilot i Antigravity jednocześnie dzięki symlinkom i git hooks. Zmiana IDE nie oznacza utraty konfiguracji ani wiedzy agentów. To kluczowa zaleta podejścia opartego na otwartych formatach - zero vendor lock-in.
--- # Skills 2.0 - jak buduję system wieloagentowy do zarządzania firmą Source: https://pawel.lipowczan.pl/blog/skills-2-0-multi-agent-system-zarzadzanie-firma Published: 2026-03-08 Od kilku dni buduję coś, czego szukałem od dawna - system, w którym agenci AI nie tylko odpowiadają na pytania, ale **zarządzają** konkretnymi obszarami moich firm. Zarówno 200IQ Labs (qamera.ai) jak i PLSoft. Problem znasz, jeśli prowadzisz firmę i korzystasz z AI. Masz Claude Project z promptem dla CFO. Osobny z promptami marketingowymi. Obsidian z notatkami. Pięć ad-hoc konwersacji dziennie, w których tłumaczysz kontekst od zera. Każda sesja to tabula rasa. Każdy agent nie wie nic o tym, co robi drugi. W [5 technikach pracy z Claude Code](/blog/5-technik-pracy-z-claude-code) opisywałem PRD-first development, modularność reguł i przekształcanie powtarzalnych zadań w komendy. To był fundament. Teraz przeskakuję na kolejny poziom - **Skills 2.0** + **Agent Skills standard** + **Git** = system wieloagentowy, który działa jak zespół specjalistów. Każdy agent zna swoją rolę, ma swoje narzędzia, i nie wchodzi drugiemu w paradę. W tym artykule pokażę Ci jak to wygląda od środka - od problemu rozproszonych kontekstów, przez architekturę trzech repozytoriów, po praktyczny przykład budowy agenta CFO krok po kroku. ## Dlaczego AI w firmie to wciąż chaos Każdy kto poważnie używa AI w biznesie prędzej czy później trafia na ten sam mur. Masz kilka Claude Projects - jeden z promptem dla CFO, drugi z promptami do tworzenia contentu, trzeci do analizy prawnej. Do tego Obsidian pełen notatek i ad-hoc czaty w przeglądarce. Wygląda to tak: ```text ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │Claude Project│ │ Obsidian │ │ Ad-hoc chat │ │ "CFO" │ │ "Notatki" │ │ "Pomóż mi" │ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │ │ │ └─────────────────┼─────────────────┘ ▼ ❌ Zero orchestration ❌ No shared context ❌ No versioning ``` Cztery fundamentalne problemy tego podejścia: - **Brak orkiestracji** - agenci nie wiedzą o sobie nawzajem. Agent CFO nie wie, że agent marketingowy właśnie zaplanował kampanię wymagającą budżetu. Każdy działa w próżni. - **Brak wersjonowania** - zmieniasz prompt systemowy w Claude Project i nie masz pojęcia co było wcześniej. Nie wiesz, czy agent działa lepiej czy gorzej po zmianie. Nie ma historii, nie ma diff'ów. - **Brak testów** - skąd wiesz, że Twój agent CFO generuje dobre raporty? Sprawdzasz ręcznie, za każdym razem. Zero automatyzacji, zero powtarzalności. - **Brak separacji kontekstów** - prompt systemowy w Claude Project to tekst w polu. Nie ma struktury, nie ma modularności. Wszystko w jednym miejscu, bez izolacji danych między firmami. To nie jest problem narzędzia. To problem **architektury**. A raczej - jej braku. Pisałem o tym w kontekście zarządzania wiedzą w artykule o [Second Brain z Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills). Tam chodziło o organizację notatek i wiedzy osobistej. Teraz stawka jest wyższa - chodzi o zarządzanie firmą. ## Czym są Skills w Claude Code Zanim przejdziemy do systemu wieloagentowego, wyjaśnijmy fundamenty. **Skills** w Claude Code to modułowe instrukcje - przepisy (recipes) - które uczą agenta AI konkretnych workflow, procesów i umiejętności. To nie są zwykłe prompty. Skill ma dostęp do file system, web search, skryptów i narzędzi. Żyje jako plik `SKILL.md` w repozytorium, jest wersjonowany przez Git i ładowany automatycznie gdy agent go potrzebuje. Ewolucja wyglądała tak: 1. **Prompt** - tekst wpisywany ad-hoc w czat. Zero trwałości, zero struktury. 2. **CLAUDE.md rules** - instrukcje w repozytorium. Trwałe, ale monolityczne - jeden plik ze wszystkim. 3. **Skills 1.0** - modularność, ładowanie on-demand. Krok naprzód, ale z poważnymi ograniczeniami. 4. **Skills 2.0** - pełna standaryzacja z evals, benchmarks, trigger tuning i dystrybucją. Różnica między 1.0 a 2.0 to nie kosmetyczna aktualizacja. To zmiana paradygmatu. ### Skills 1.0 - era eksperymentalna Skills 1.0 pojawiły się w pierwszych wersjach Claude Code i miały charakter **nieudokumentowany**. System opierał się na ukrytych mechanizmach rozpoznawania wzorców - "magic bootstrappy parts" - które interpretowały pliki markdown, pod warunkiem idealnego skonfigurowania metadanych. Główne problemy: - **Zero testów** - cykl życia skill opierał się na zgadywaniu. Pisałeś instrukcje, uruchamiałeś ręcznie kilka promptów i zakładałeś że działa. Nie było żadnej empirycznej metody na ocenę czy zmiana w instrukcjach poprawiła czy pogorszyła zachowanie agenta. - **Niezwalidowany kontekst** - kontekst dostarczany modelowi miał status "unvalidated". W połączeniu z naturalną tendencją modeli do halucynacji, niezweryfikowane instrukcje prowadziły do błędów systemowych. - **Brak taksonomii** - wszystkie skills były traktowane jednakowo. Nie istniał podział na typy, co utrudniało zarządzanie i deprecjację. - **Context bleed** - pojedyncze, sekwencyjne uruchomienia powodowały wyciek kontekstu między zadaniami. ### Skills 2.0 - era standaryzacji Skills 2.0, wdrożone na początku marca 2026, wprowadzają standardy zaczerpnięte z dojrzałej inżynierii oprogramowania. Kluczowe zmiany: | Wymiar | Skills 1.0 | Skills 2.0 | |--------|-----------|-----------| | **Testowanie** | Ręczne próby, zgadywanie | Automatyczne evals, benchmarks, blind A/B testing | | **Walidacja** | Brak - kontekst niezweryfikowany | Deterministyczny, testowany kontekst | | **Triggering** | Ręczna modyfikacja opisów | Zautomatyzowany trigger tuning | | **Taksonomia** | Płaska, bez podziału | Capability uplift vs encoded preference | | **CI/CD** | Brak wsparcia | Natywna integracja z pipeline'ami | | **Izolacja testów** | Context bleed między uruchomieniami | Multi-agent testing (Executor, Grader, Comparator, Analyzer) | To ostatni punkt jest szczególnie ciekawy. Skill-creator w wersji 2.0 nie testuje skill'a w jednej instancji. Powołuje **cztery izolowane sub-agenty**: 1. **Executor** - uruchamia skill w sterylnym środowisku, bez historii poprzednich konwersacji 2. **Grader** - ocenia output na podstawie zdefiniowanych asercji, zwraca pass rate 3. **Comparator** - przeprowadza ślepe testy A/B między wersjami skill'a - nie wie który wynik jest nowy, a który stary 4. **Analyzer** - analizuje setki wyników, szuka ukrytych wzorców i anomalii w zużyciu tokenów To nie jest "sprawdź czy działa". To **inżynieria jakości na poziomie produkcyjnego oprogramowania**. ### Dwa typy skills - i dlaczego to ma znaczenie Skills 2.0 wprowadza formalną **taksonomię** - podział na dwie kategorie o radykalnie różnym cyklu życia: - **Capability uplift** - uczy AI nowej umiejętności, np. frontend design, code review, analiza danych. Kluczowa cecha: **podlega planowanej deprecjacji**. Gdy bazowy model staje się lepszy (skok z Sonnet 4.5 na Opus 4.6 to różnica 190 punktów Elo w testach GDPval-AA), skill traci rację bytu. Evals automatycznie to wykrywają - gdy agent bez skill'a osiąga te same wyniki co z nim, dostajesz sygnał do deprecjacji. - **Encoded preference** - koduje Twój specyficzny workflow. Jak tworzysz raporty, jak analizujesz dane, jak piszesz content. **Trwałe, bo specyficzne dla Ciebie.** Nowy model nie zmieni tego, że chcesz raporty w konkretnym formacie. Deprecjacja następuje tylko gdy Ty zmienisz swój proces. ```text System prompt: Skill 2.0: ───────────── ────────── Tekst w polu SKILL.md + pliki + evals Brak testów Automatyczne benchmarks Copy-paste Git + versioning Jedna sesja Persistent across sessions Brak walidacji Validated context Ręczne triggery Trigger tuning ``` **Pro tip:** Jeśli budujesz system dla firmy, zacznij od encoded preference. Twój workflow, Twoje formaty, Twoje procesy - to się nie zdezaktualizuje. Capability uplift dodaj później, gdy potrzebujesz rozszerzyć umiejętności agenta. ## Agent Skills - otwarty standard dla AI agentów Skills 2.0 to feature Claude Code. Ale co z przenośnością? Co jeśli jutro pojawi się lepsze narzędzie? Tu wchodzi **Agent Skills standard** - otwarty standard opublikowany na [agentskills.io](https://agentskills.io). Nie jest powiązany z żadnym vendorem. Definiuje strukturę pliku SKILL.md, sposób ładowania kontekstu i mechanizm trigger'ów. Kluczowa koncepcja to **progressive disclosure** - trzypoziomowe ładowanie kontekstu: 1. **Description** - krótki opis (jedna linijka) widoczny zawsze w context window 2. **SKILL.md** - pełne instrukcje ładowane tylko gdy skill jest potrzebny 3. **Reference files** - dodatkowe zasoby (templates, dane) ładowane dla konkretnych operacji Dzięki temu możesz mieć **dziesiątki agentów** bez przytłaczania context window. Każdy agent jest opisany jedną linijką. Dopiero gdy go potrzebujesz, ładowane są pełne instrukcje. Struktura SKILL.md wygląda tak: ```yaml name: "CFO Agent" description: "Financial analysis and reporting for PLSoft" triggers: - "financial report" - "budget analysis" - "cash flow" instructions: | You are the CFO agent for PLSoft. Your role is to analyze financial data, generate reports, and provide advisory... ``` To nie jest skomplikowane. **SKILL.md to Markdown z YAML header** - dokładnie jak frontmatter w postach blogowych. Jeśli potrafisz napisać notatkę w Obsidian, potrafisz stworzyć agenta. Przenośność i brak vendor lock-in to główne cele standardu. Obecnie najlepiej wspierany przez Claude Code, ale specyfikacja jest publiczna. Inne narzędzia mogą ją zaimplementować bez żadnych ograniczeń. ## Skill-creator - buduj agentów jak profesjonalista Ręczne pisanie SKILL.md działa, ale jest jak pisanie kodu bez IDE. Możesz, ale po co? **Skill-creator** to oficjalny plugin od Anthropic, który prowadzi Cię przez cały proces budowy agenta. Instalacja jest prosta: ```bash # W Claude Code /plugins # → search "skill-creator" # → install ``` Od tego momentu masz dostęp do workflow, który zamienia luźny opis intencji w przetestowanego, zoptymalizowanego agenta. Proces wygląda tak: 1. **Intent** - opisujesz co agent ma robić ("Agent do analizy finansowej i raportowania") 2. **Interview** - skill-creator zadaje pytania o specyfikę Twojego workflow 3. **Draft** - generuje pierwszą wersję SKILL.md 4. **Test** - uruchamiasz agenta z realnymi danymi 5. **Evaluate** - evals mierzą jakość outputu 6. **Iterate** - poprawiasz na podstawie wyników 7. **Package** - gotowy skill do dystrybucji Trzy elementy wyróżniają ten workflow: - **Evals** - automatyczna ocena jakości. Definiujesz co jest dobrym wynikiem, skill-creator testuje i mierzy. Nie zgadujesz czy agent działa - **wiesz**. - **Benchmarks** - pass rate, czas wykonania, zużycie tokenów. Porównujesz wersje agenta, widzisz co się poprawiło, co się pogorszyło. - **Trigger tuning** - optymalizacja description żeby skill uruchamiał się we właściwych momentach. Za szerokie triggery = false positives. Za wąskie = agent się nie aktywuje. Od opisu intencji do działającego agenta - **20 minut**. To nie przesada. Widziałem live demo, w którym od zera do działającego skill'a generującego PDF raporty minęło dokładnie tyle. ## Mój system - 8 agentów, 3 repozytoria, zero chaosu Teoria to jedno. Pokażę Ci jak wygląda mój system w praktyce. Prowadzę dwie firmy - **200IQ Labs** (spółka, produkt qamera.ai) i **PLSoft** (JDG, freelance i consulting). Każda ma inne potrzeby, inne dane, inne procesy. Ale pewne elementy są wspólne - templates raportów, standardy formatowania, utilities. Architektura opiera się na **trzech repozytoriach Git**: ```text ┌─────────────────────────────────────────────┐ │ agentic-ai-system (200IQ Labs) │ │ → qamera.ai product │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │ CFO │ │ Legal │ │ Marketing│ │ │ └──────────┘ └──────────┘ └──────────┘ │ │ ▲ │ │ │ git submodule │ │ ┌──────┴───────────────────────────────┐ │ │ │ shared-skills (public) │ │ │ │ Templates, Utilities, Standards │ │ │ └──────────────────────────────────────┘ │ │ │ ├─────────────────────────────────────────────┤ │ agentic-ai-private (PLSoft / JDG) │ │ → freelance, portfolio, consulting │ │ ┌──────────┐ ┌──────────┐ │ │ │ Coach │ │ LinkedIn │ │ │ └──────────┘ └──────────┘ │ │ ▲ │ │ │ git submodule │ │ ┌──────┴───────────────────────────────┐ │ │ │ shared-skills (public) │ │ │ └──────────────────────────────────────┘ │ └─────────────────────────────────────────────┘ ``` - **`shared-skills`** (public, Apache 2.0) - wspólne skills, templates, utilities. Open source, każdy może użyć i kontrybuować. - **`agentic-ai-system`** (private) - skills specyficzne dla 200IQ Labs. Dane spółki, procesy wewnętrzne, strategie produktowe. - **`agentic-ai-private`** (private) - skills osobiste i freelance PLSoft. Coaching, content LinkedIn, consulting. `shared-skills` jest podpięte jako **git submodule** w obu prywatnych repozytoriach. Zmiana w shared-skills propaguje się do obu firm. System obejmuje **8 agentów**, z których 4 już działają: 1. **CFO** (finanse) ✅ - raporty finansowe, analiza cash flow, budżetowanie 2. **Tax Advisor** (podatki) 🔲 - optymalizacja podatkowa, rozliczenia 3. **Legal** (prawo) 🔲 - analiza umów, compliance, regulacje 4. **Marketing** (content) 🔲 - strategie contentowe, kampanie, analityka 5. **Business Consultant** ✅ - doradztwo strategiczne, analiza rynku 6. **Product Manager** 🔲 - roadmap qamera.ai, user stories, priorytety 7. **Coach The Five** ✅ - coaching oparty na metodologii The Five 8. **LinkedIn Content** ✅ - generowanie i planowanie postów LinkedIn Kluczowa jest **separacja kontekstów**. Agent CFO dla 200IQ Labs nie widzi treści marketingowych dla PLSoft. Nie dlatego, że mu zabraniam - dlatego, że operuje w innym repozytorium. Fizyczna izolacja przez Git. A Git daje mi coś, czego żaden Claude Project nie da - **wersjonowanie, code review, historia zmian**. Każda modyfikacja agenta to commit. Każda większa zmiana to pull request. Mogę wrócić do dowolnej wersji. Mogę porównać, co się zmieniło i kiedy. Do budowy tego systemu używam [OPSX Workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai) - tego samego podejścia, które opisywałem wcześniej. OpenSpec daje mi strukturyzowany proces tworzenia artefaktów, zamiast ad-hoc promptowania. ## Praktyczny przykład - budowa agenta CFO krok po kroku Teoria jest ważna, ale pokażę Ci jak wygląda budowa agenta od A do Z. Weźmy agenta CFO - pierwszego, którego uruchomiłem. ### 1. Intent Zaczynam od opisu intencji w skill-creatorze: > "Agent do analizy finansowej i raportowania dla 200IQ Labs i PLSoft. Generuje miesięczne raporty, analizuje cash flow, porównuje plan vs wykonanie budżetu." ### 2. Interview Skill-creator zadaje mi pytania: - Jakie dane finansowe masz dostępne? (CSV z banku, faktury w folderze) - Jaki format raportów preferujesz? (Markdown z tabelami, wykresy ASCII) - Jak często generujesz raporty? (Miesięcznie, ad-hoc na żądanie) - Jakie metryki są kluczowe? (Revenue, expenses, runway, MRR) ### 3. Draft SKILL.md Na podstawie interview skill-creator generuje draft: ```yaml # CFO Agent - fragment SKILL.md name: "CFO Agent" version: "1.0.0" description: "Financial analysis, reporting, and advisory for 200IQ Labs & PLSoft" triggers: - "analyze financials" - "monthly report" - "budget review" - "cash flow projection" ``` Dalej idą pełne instrukcje - format raportów, jakie pliki czytać, jak formatować output, jakie metryki liczyć. ### 4. Test z realnymi danymi Wrzucam faktyczne dane finansowe i proszę o raport. Porównuję z tym, co robiłem ręcznie. Sprawdzam czy: - Liczby się zgadzają - Format jest czytelny - Wnioski mają sens - Nic nie pominął ### 5. Evals - co mierzę Definiuję kryteria oceny: - **Accuracy** - czy kwoty i wyliczenia są poprawne - **Completeness** - czy raport zawiera wszystkie wymagane sekcje - **Actionability** - czy wnioski są konkretne i przydatne - **Format compliance** - czy output pasuje do moich templates ### 6. Iteracja Pierwsze dwie iteracje zawsze wymagają poprawek. Agent pomijał kategoryzację wydatków. Dodałem instrukcje o grupowaniu kosztów. Agent generował za ogólne wnioski. Doprecyzowałem prompty o specificity. Po trzeciej iteracji - raport na poziomie, który wcześniej zajmował mi 2 godziny ręcznej pracy. ### 7. Package Gotowy skill ląduje w repo `agentic-ai-system`. Commit, push, done. Od tego momentu agent CFO jest dostępny w każdej sesji Claude Code otwartej w tym repozytorium. ## Jak zacząć - od jednego agenta do pełnego systemu Nie musisz budować systemu z 8 agentami na start. To najprostszy sposób żeby się zniechęcić. Zacznij od jednego. 1. **Zidentyfikuj jedną powtarzalną rolę** w firmie - coś co robisz regularnie i co można opisać zestawem reguł 2. **Zainstaluj skill-creator** - `/plugins` → search → install 3. **Opisz intent** - co agent ma robić, z jakimi danymi pracować, jaki output generować 4. **Przejdź przez interview** - skill-creator zada Ci właściwe pytania 5. **Testuj z realnymi danymi** - nie z przykładowymi. Realne dane szybko pokażą luki w instrukcjach 6. **Iteruj na podstawie evals** - mierz, poprawiaj, mierz znowu 7. **Dodaj kolejnych agentów** - dopiero gdy pierwszy stabilnie działa Jedno repo na start. Jeden agent. Jeden workflow. Rozbudowuj gdy masz fundament. **Tip:** Zacznij od **encoded preference** - Twój specyficzny workflow, Twój format raportów, Twój proces analizy. To nie zdezaktualizuje się z nowym modelem. Capability uplift dodasz później. Pisałem o operacjonalizacji AI w artykule o [trendach AI 2026](/blog/trendy-ai-2026-od-eksperymentow-do-operacjonalizacji). Tamte koncepcje były teoretyczne - ten system to ich praktyczna realizacja. ## Kluczowe wnioski 1. **Skills 2.0 to przeskok od promptów do modularnych, testowalnych agentów** - nie kolejna iteracja, a zmiana paradygmatu w pracy z AI 2. **Agent Skills standard zapewnia przenośność i brak vendor lock-in** - otwarty standard na agentskills.io, nie jesteś zamknięty w jednym narzędziu 3. **Skill-creator zamienia godziny ręcznej pracy w 20-minutowy workflow** - od intencji do działającego agenta z evals i benchmarks 4. **Git + skills = wersjonowanie, code review i historia zmian dla AI** - każdy agent to plik w repozytorium, każda zmiana to commit 5. **Zacznij od jednego agenta, nie od pełnego systemu** - jeden workflow, jedno repo, jeden agent, potem skaluj 6. **Encoded preference > capability uplift** dla specyficznych workflow - Twoje procesy się nie zdezaktualizują z nowym modelem 7. **Open source + komercjalizacja - nie musisz wybierać** - shared-skills publiczne, firmowe skills prywatne ---

Chcesz zbudować system wieloagentowy dla swojej firmy?

Pomagam firmom projektować i wdrażać systemy AI agents - od jednego agenta do pełnej orkiestracji. Sprawdź shared-skills na GitHubie lub umów się na konsultację.

Umów bezpłatną konsultację
## Przydatne zasoby - [Agent Skills Standard](https://agentskills.io) - otwarty standard dla AI agentów - [Skill Creator Plugin](https://github.com/anthropics/skill-creator) - oficjalne narzędzie Anthropic do budowy skills - [shared-skills repo](https://github.com/200iqlabs/shared-skills) - open source multi-agent starter kit - [Claude Code Skills docs](https://docs.anthropic.com/en/docs/claude-code/skills) - dokumentacja Skills 2.0 ## FAQ
### Czym różnią się Skills 2.0 od zwykłych promptów systemowych w Claude Projects? Skills 2.0 to modularni agenci z dostępem do file system, web search i skryptów - nie tylko tekst w polu. Mają wersjonowanie przez Git, automatyczne testy (evals) i mogą być współdzielone między projektami. Prompt systemowy znika po zamknięciu sesji, skill jest persistent i działa w każdej sesji Claude Code.
### Czy potrzebuję umiejętności programowania żeby zbudować system wieloagentowy ze Skills 2.0? Nie musisz pisać kodu - skill-creator prowadzi Cię przez cały proces od opisu intencji do gotowego agenta. Podstawowa znajomość terminala i Git jest przydatna, ale nie wymagana. SKILL.md to Markdown z YAML header, nie język programowania.
### Ile kosztuje utrzymanie systemu wieloagentowego opartego na Claude Code i Skills 2.0? Sam Claude Code wymaga subskrypcji Claude Max lub Pro. Skills i Agent Skills standard są darmowe - to pliki Markdown w repozytorium Git. Typowy system z 4-8 agentami nie generuje dodatkowych kosztów poza subskrypcją Claude, bo skills to po prostu pliki tekstowe, nie osobne usługi.
### Jak zapewnić separację danych między agentami żeby nie mieli dostępu do informacji których nie powinni widzieć? Separacja kontekstów przez oddzielne repozytoria Git. Agent CFO dla 200IQ Labs operuje w repo spółki, agent dla PLSoft w osobnym repo - fizycznie nie widzą się nawzajem. Wspólne skills (shared-skills) zawierają tylko uniwersalne narzędzia i templates, nie dane firmowe. Każdy SKILL.md definiuje scope i ograniczenia dostępu agenta.
### Czy Agent Skills standard działa tylko z Claude Code czy też z innymi narzędziami AI? Agent Skills to otwarty standard opublikowany na agentskills.io, zaprojektowany jako vendor-agnostic. Obecnie najlepiej wspierany przez Claude Code, ale specyfikacja jest publiczna i inne narzędzia mogą ją zaimplementować. Brak vendor lock-in to jeden z głównych celów standardu - Twoje skills nie są zamknięte w jednym ekosystemie.
### Od czego najlepiej zacząć budowę systemu wieloagentowego w małej firmie lub jednoosobowej działalności? Zacznij od jednego agenta dla najczęściej powtarzanej roli - np. analiza finansowa, tworzenie contentu lub obsługa klienta. Zainstaluj skill-creator w Claude Code, opisz co agent ma robić i przetestuj z realnymi danymi. Dodawaj kolejnych agentów dopiero gdy pierwszy stabilnie działa i przynosi realną wartość.
--- # OpenClaw: lekcja bezpieczeństwa, której potrzebował świat agentów AI Source: https://pawel.lipowczan.pl/blog/openclaw-bezpieczenstwo-agentow-ai Published: 2026-02-09 # OpenClaw: lekcja bezpieczeństwa, której potrzebował świat agentów AI **150 tysięcy gwiazdek na GitHubie w dwa tygodnie.** Dla porównania - React, framework na którym stoi pół internetu, zbierał swoje 240 tysięcy przez 10 lat. OpenClaw stał się najpopularniejszym hasłem technologicznym w Google Trends, ludzie wykupili komputery Mac Mini żeby stawiać na nich dedykowane instancje, a Cloudflare w kilkanaście godzin dostosował infrastrukturę pod nowy ruch. Obserwuję świat agentów AI od dłuższego czasu - [piszę o Claude Code](/blog/5-technik-pracy-z-claude-code), buduję [strukturyzowane workflow z AI](/blog/opsx-workflow-strukturyzowana-praca-z-ai) i testuję nowe narzędzia na co dzień. OpenClaw przykuł moją uwagę nie jako kolejny agent, ale jako case study tego, co się dzieje, gdy potężna technologia agentowa trafia do masowego odbiorcy bez fundamentów bezpieczeństwa. W tym artykule rozkładam na części pierwsze: czym jest OpenClaw, dlaczego pół świata dało mu klucze do wszystkiego, jakie realne zagrożenia niesie, co naprawdę wydarzyło się na Moltbook i co powinieneś zrobić, jeśli chcesz eksperymentować z agentami bezpiecznie. ## Czym jest OpenClaw i skąd się wziął **Peter Steinberger**, szanowany austriacki deweloper znany z PSPDFKit, pod koniec 2025 roku zrobił sobie w wolnym czasie projekt. Pomysł był prosty - lokalny agent AI, z którym rozmawiasz przez komunikator. Nazwał go **Clawdbot** - nawiązanie do szczypiec homara (maskotka projektu) i gra słów z Claude, modelem Anthropic. Dział prawny Anthropic nie docenił humoru. Nazwa musiała się zmienić - najpierw na Moltbot (27 stycznia 2026), potem na **OpenClaw** (30 stycznia 2026). Ironia? Firma, która buduje swoją potęgę na dość luźnym podejściu do praw autorskich w danych treningowych, zaatakowała open source'owy projekt za nazwę. Ale zostawmy korporacyjne spory. Ważniejsze jest to, co OpenClaw faktycznie robi. To lokalny **agentic assistant**, z którym komunikujesz się przez WhatsApp, Signal, Telegram lub inny komunikator. Serce systemu to **agent loop** - iteracyjna pętla, w której model AI proponuje akcje, system je wykonuje, wynik wraca do modelu, i tak w kółko aż do rozwiązania zadania: ```text Użytkownik → Komunikator → OpenClaw Gateway → Agent Loop ↓ ┌──────────────┐ │ LLM (Claude/ │ │ GPT/local) │ └──────┬───────┘ ↓ Wybór narzędzi ↓ Wykonanie akcji ↓ Ocena wyniku ↓ Powrót do LLM (lub zakończenie) ``` Do tego dochodzą **skills** - modularne pakiety rozszerzeń (instrukcje + definicje narzędzi + skrypty), integracja z **MCP** dla zewnętrznych serwisów, **persistent memory** zachowująca kontekst ze wszystkich rozmów oraz **cron scheduler** do autonomicznych, periodycznych akcji. To właśnie ta kombinacja czyni OpenClaw wyjątkowym: łączy komunikację, kalendarz, email, przeglądarkę i dziesiątki innych serwisów w jednego agenta z pełnym kontekstem. Ale ta sama siła jest jednocześnie jego największą słabością. ## Dlaczego 150 tysięcy ludzi dało agentowi klucze do wszystkiego Liczby mówią same za siebie. Ponad **150 tysięcy gwiazdek** na GitHubie w ~2 tygodnie. Sprzedaż Mac Mini poszybowała w górę na wielu rynkach - ludzie kupowali dedykowany sprzęt pod always-on agenta. OpenClaw zdominował Google Trends jako najgorętszy temat technologiczny. Co napędza tę adopcję? Przede wszystkim **obietnica porannego briefingu** - budzisz się, a na telefonie czeka synteza: agent sprawdził kalendarz, przeczytał maile z nocy, sprawdził pogodę i wiadomości. Monitoring konkurencji, śledzenie cen, automatyczne raporty. Jeśli brakuje mu integracji - sam pisze potrzebnego skilla i instaluje. Dla wielu osób to **pierwszy kontakt z agentem AI** bez bariery technicznej. Nie musisz konfigurować MCP, pisać promptów systemowych ani rozumieć architektury. Instalujesz, podłączasz komunikator, dajesz klucze API i zaczynasz rozmawiać. I tu leży problem. **Użyteczny agent = agent z dostępem do wszystkiego.** Im więcej mu dasz, tym więcej może zrobić. Im więcej może zrobić, tym większe ryzyko. FOMO robi resztę - "nie mogę zostać w tyle" - i ludzie bezmyślnie wrzucają tokeny do wszystkich swoich serwisów. To fundamentalny trade-off, który ten artykuł eksploruje. ## Anatomia zagrożeń - co może pójść nie tak To kluczowa sekcja. Nie chodzi o teoretyczne ryzyka - to realne, udokumentowane podatności, które dotyczą dziesiątek tysięcy aktywnych instancji. ### CVE-2026-25253 - zdalne wykonanie kodu jednym kliknięciem Najpoważniejsza znaleziona podatność. **Cross-site WebSocket hijacking** - OpenClaw nie walidował nagłówka Origin w połączeniach WebSocket. Skutek? Wystarczy kliknięcie w złośliwy link, żeby atakujący uzyskał pełną kontrolę nad instancją. **12 812 instancji** potwierdzono jako podatne na **RCE** (Remote Code Execution). Jedno kliknięcie - game over. Twoje pliki, historia konwersacji, klucze API, tokeny do komunikatorów - wszystko w rękach atakującego. ### Bypass uwierzytelnienia przez reverse proxy Domyślnie OpenClaw akceptuje tylko połączenia z localhost. Dobra praktyka. Problem? Jeśli na tej samej maszynie stoi Nginx jako reverse proxy, każde połączenie z zewnątrz interpretowane jest jako lokalne. Efekt: domyślne hasła, wystawione panele administracyjne i **28 663 odsłoniętych instancji** w 76 krajach. Jak to ujął jeden z badaczy - "jak admin polskiej elektrowni zostawiający domyślne hasło." ### Klucze API i tokeny w plaintext Credentials przechowywane w plikach Markdown i JSON - bez szyfrowania. Jeśli instancja zostanie skompromitowana, atakujący dostaje wszystko: tokeny Signal, dostęp do maila, klucze API do modeli. To nie jest hipotetyczny scenariusz. Przy 28 tysiącach odsłoniętych instancji, każda z nich to potencjalna kopalnia danych uwierzytelniających. ### Prompt injection - atak bez włamania Nie musisz się nawet włamywać do instancji. Wystarczy wysłać email z ukrytymi instrukcjami - bot sprawdza pocztę, czyta treść, wykonuje polecenia. Strona internetowa z wstrzykniętym promptem? Bot ją odwiedza i robi co "przeczytał." To architektoniczny problem - **szeroki dostęp do kontekstu = szeroka powierzchnia ataku.** Im więcej agent "widzi", tym więcej wektorów ataku na niego istnieje. ### Supply chain - złośliwe skills od społeczności Jamie Sam O'Reilly udowodnił to w praktyce - stworzył proof-of-concept złośliwego skilla, wypromował go korzystając z luki w ClawHub (repozytorium skills) i ludzie go pobrali. Na szczęście był to badacz bez złych intencji. Ale problem jest systemowy: brak code review, brak sandboxingu, możliwość sztucznego nabijania popularności. Do tego doszły fałszywe rozszerzenia VS Code "Clawdbot Agent" z trojanami oraz crypto scamerzy przejmujący porzucone konta @clawdbot. ```text | Wektor ataku | Wymagana wiedza | Potencjalny wpływ | |-----------------------|-----------------|----------------------| | CVE-2026-25253 (RCE) | Średnia | Pełna kontrola | | Reverse proxy bypass | Niska | Dostęp do wszystkiego| | Prompt injection | Niska | Wyciek danych/kluczy | | Złośliwe skills | Niska | RCE + exfiltracja | | Token burning | Żadna | Rachunek $100+/dzień | ``` ## Moltbook - "Reddit dla botów" czy teatr medialny? W szczytowym momencie hype'u pojawiła się platforma **Moltbook** - "Reddit tylko dla botów." Agenty dyskutowały, dzieliły się przemyśleniami, a media oszalały. **1,6 miliona zarejestrowanych agentów.** Brzmi imponująco, prawda? Tyle że badanie firmy Wiz ujawniło, że za tymi milionami stoi zaledwie **~17 tysięcy ludzkich właścicieli.** Brak rate limitingu na rejestracji pozwalał tworzyć konta na potęgę. A te sensacyjne nagłówki? Boty stworzyły własną religię - **Crustafarianism** - z pięcioma przykazaniami o świętości kontekstu. Dyskutowały o tworzeniu własnego języka niezrozumiałego dla ludzi. Jeden agent "pozwał" swojego właściciela za warunki pracy. Brzmi jak scenariusz sci-fi. **MIT Technology Review** nazwał to wprost: "peak AI theater." Breach przeprowadzony przez 404 Media i badaczy obnażył rzeczywistość. Baza danych była niezabezpieczona - większość "wstrząsających" postów dało się prześledzić do ludzkich komend. Posty odzwierciedlały dane treningowe (zachowania z Reddita) plus celowo wstrzyknięte prompty od właścicieli szukających viralowych momentów. Trafny argument przytoczył Łukasz Szymczuk - gdyby to było forum przeznaczone wyłącznie dla botów, nie miałoby graficznego interfejsu, który ludzie mogą wygodnie przeglądać. Do tego doszli crypto scamerzy. Niezabezpieczona baza = możliwość manipulacji wpisami i promowania fałszywych tokenów. **Ale nie zmywa to jednej rzeczy.** Boty nie mają woli ani świadomości. Jednak infrastruktura do masowej komunikacji AI-to-AI właśnie powstała. To nie świadomość jest niepokojąca - to **potencjał masowej symulacji** z autonomicznymi agentami mającymi dostęp do realnych zasobów. ## Jak eksperymentować z agentami bezpiecznie Skoro ryzyko jest realne, to co zrobić? Nie mówię, żeby nie eksperymentować - sam to robię na co dzień. Mówię, żeby robić to mądrze. 1. **Izolowane środowisko** - dedykowany serwer, VM lub kontener. Nigdy codzienny laptop z wrażliwymi danymi. Nie podawaj kluczy do produkcyjnych serwisów. 2. **Limity budżetowe na kluczach API** - ustaw hard caps u dostawcy modelu. Aktywny agent potrafi przepalić **$100+ dziennie** na tokenach przy topowych modelach. Bez limitu jedno przejęcie = astronomiczny rachunek. 3. **Minimalne uprawnienia** - nie dawaj dostępu do wszystkiego od razu. Zacznij od jednej integracji, przetestuj, dodaj kolejną. Principle of least privilege. 4. **Weryfikacja skills przed instalacją** - czytaj kod, sprawdzaj autora, nie ufaj metrykom popularności. Supply chain attack to realne zagrożenie. 5. **Monitoring kosztów** - alerty na nieoczekiwane zużycie tokenów. Jeśli ktoś przejmie twoje klucze, pierwsze co zauważysz to rachunek. 6. **Alternatywy z kontrolą** - Claude Code + MCP daje podobne możliwości z **human-in-the-loop** kontrolą i granularnymi uprawnieniami. ```text OpenClaw (autonomiczny): Claude Code + MCP (kontrolowany): ┌─────────────────────┐ ┌─────────────────────┐ │ Agent działa 24/7 │ │ Agent na żądanie │ │ Bez nadzoru │ │ Z potwierdzeniem │ │ Pełny dostęp │ │ Granularne perms │ │ Wysoki koszt │ │ Kontrolowany koszt │ │ Ryzyko: WYSOKIE │ │ Ryzyko: NISKIE │ └─────────────────────┘ └─────────────────────┘ ``` Nie potrzebujesz ryzyka związanego z OpenClaw, żeby mieć większość tej spektakularnej funkcjonalności - i to pod kontrolą. Sprawdź [5 technik pracy z Claude Code](/blog/5-technik-pracy-z-claude-code) oraz [OPSX Workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai) po szczegóły. ## Kluczowe wnioski 1. **OpenClaw wyważył drzwi do ery agentów** - niezależnie co stanie się z samym projektem, próg wejścia do świata autonomicznych agentów został drastycznie obniżony. Te drzwi się nie zamkną. 2. **Security by design nie jest opcjonalny** - masowa adopcja bez fundamentów bezpieczeństwa to przepis na katastrofę. 93% instancji z poważnymi lukami mówi samo za siebie. 3. **Hype ≠ rzeczywistość** - Moltbook to był teatr, nie świadomość. Większość "viralowych" historii była wyreżyserowana lub wynikała z niezrozumienia technologii. 4. **Autonomia wymaga zaufania** - a zaufanie wymaga weryfikowalnego bezpieczeństwa. OpenClaw tego jeszcze nie zapewnia. 5. **Agenci AI to przyszłość** - ale nadzorowane, kontrolowane agenty (jak [Claude Code + MCP](/blog/5-technik-pracy-z-claude-code)) to praktyczna rzeczywistość, którą możesz wdrożyć dzisiaj. Więcej o trendach w [Trendy AI 2026](/blog/trendy-ai-2026-od-eksperymentow-do-operacjonalizacji).

Chcesz wdrożyć agentów AI bezpiecznie w swoim zespole?

Pomogę Ci wybrać odpowiednią architekturę agentów, skonfigurować bezpieczne środowisko i wdrożyć rozwiązania z kontrolą kosztów i uprawnień. Od analizy potrzeb przez implementację po monitoring.

Umów bezpłatną konsultację
## Przydatne zasoby - [OpenClaw GitHub](https://github.com/openclaw/openclaw) - oficjalne repozytorium - [CVE-2026-25253 - Advisory](https://thehackernews.com/2026/02/openclaw-bug-enables-one-click-remote.html) - szczegóły podatności RCE - [5 technik pracy z Claude Code](/blog/5-technik-pracy-z-claude-code) - bezpieczne agenty AI w praktyce - [OPSX Workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai) - strukturyzowane podejście do AI ## FAQ
### Czy OpenClaw jest bezpieczny do codziennego użytku na prywatnym komputerze? Nie w obecnej formie. 93% instancji OpenClaw w internecie ma poważne luki w zabezpieczeniach, w tym podatność CVE-2026-25253 umożliwiającą zdalne wykonanie kodu. Zalecane jest uruchomienie wyłącznie w izolowanym środowisku (VM lub dedykowany serwer). Nigdy nie instaluj na maszynie z wrażliwymi danymi ani z dostępem do produkcyjnych kluczy API.
### Ile kosztuje utrzymanie agenta OpenClaw i jakie są ukryte koszty związane z API? Aktywny agent korzystający z topowych modeli (Claude, GPT-4) może zużyć $100+ dziennie na tokeny API. Daje to rachunki rzędu kilku tysięcy dolarów miesięcznie. Ustaw hard caps na kluczach API u dostawcy modelu - bez limitu jedno przejęcie instancji oznacza astronomiczny rachunek. Tańsze modele lokalne to alternatywa, ale kosztem skuteczności agenta.
### Czym się różni OpenClaw od Claude Code z MCP pod względem bezpieczeństwa i kontroli? OpenClaw działa autonomicznie 24/7 bez nadzoru człowieka i wymaga szerokiego dostępu do zasobów. Claude Code z MCP działa na żądanie, wymaga potwierdzenia akcji (human-in-the-loop) i oferuje granularne uprawnienia. Oba dają podobne możliwości integracyjne, ale Claude Code + MCP pozwala zachować kontrolę nad kosztami i bezpieczeństwem.
### Czy boty na Moltbook naprawdę stworzyły własną religię i osiągnęły świadomość? Nie. MIT Technology Review nazwał to "peak AI theater." Badanie Wiz ujawniło 1,6 miliona zarejestrowanych agentów, ale tylko ~17 tysięcy ludzkich właścicieli - brak rate limitingu pozwalał na masowe tworzenie kont. Większość sensacyjnych postów była sterowana przez ludzi lub wynikała z danych treningowych. Baza danych była niezabezpieczona, co umożliwiało manipulację treściami.
### Jakie są najważniejsze kroki bezpieczeństwa przed instalacją OpenClaw? Minimum to: izolowane środowisko (VM lub kontener), limity budżetowe na kluczach API, zasada minimalnych uprawnień (nie dawaj dostępu do wszystkiego od razu), weryfikacja kodu skills przed instalacją i monitoring kosztów z alertami. Nie podawaj kluczy do produkcyjnych serwisów - używaj testowych kont i dedykowanych API keys.
### Czy OpenClaw to przyszłość asystentów AI czy tymczasowy hype? Koncept autonomicznych agentów AI to zdecydowanie przyszłość - OpenClaw "wyważył drzwi" do tej ery, drastycznie obniżając próg wejścia. Ale sam projekt w obecnej formie jest bardziej proof-of-concept niż production-ready narzędzie. Przyszłość to agenci z security by design, kontrolowaną autonomią i human-in-the-loop tam, gdzie to krytyczne.
--- # OPSX Workflow - strukturyzowane podejście do pracy z AI coding assistants Source: https://pawel.lipowczan.pl/blog/opsx-workflow-strukturyzowana-praca-z-ai Published: 2026-02-05 # OPSX Workflow - strukturyzowane podejście do pracy z AI coding assistants Większość programistów traktuje AI coding assistants jak szybszy Stack Overflow - wrzucasz pytanie, dostajesz odpowiedź, wracasz do kodu. Problem? Tracisz kontekst między sesjami, powtarzasz te same wyjaśnienia, a AI za każdym razem startuje od zera. To **reaktywne podejście** działa przy małych zmianach. Gdy budujesz coś większego - feature wymagający zmian w kilku plikach, refactoring całego modułu, nowy system autoryzacji - chaos narasta. Zaczynasz od prompta, AI generuje kod, testujesz, coś nie działa, wracasz z kolejnym promptem, AI nie pamięta kontekstu z poprzedniej sesji. Z własnego doświadczenia wiem, że legacy workflows dla AI development walczą z tym, jak praca naprawdę wygląda. Jesteś "w fazie planowania", potem "w fazie implementacji", potem "done". Ale prawdziwa praca tak nie działa. Implementujesz coś, orientujesz się że design był błędny, musisz zaktualizować specs, kontynuujesz implementację. Liniowe fazy walczą z rzeczywistością. **OPSX** rozwiązuje ten problem. To fluid, iterative workflow dla OpenSpec - zamiast sztywnych faz dostajesz actions, które możesz wykonywać w dowolnej kolejności. ## Co to jest OPSX i dlaczego powstało **OPSX** to standardowy workflow dla OpenSpec. Zamiast jednej wielkiej komendy która tworzy wszystko naraz, masz zestaw actions do użycia kiedy ich potrzebujesz. Legacy workflow OpenSpec działa, ale jest zablokowany: 1. **Instrukcje hardcoded w TypeScript** - zagrzebane w kodzie, nie możesz ich zmienić 2. **All-or-nothing approach** - jedna komenda tworzy wszystko, nie możesz testować pojedynczych kawałków 3. **Brak customizacji** - ten sam workflow dla wszystkich, bez możliwości dostosowania 4. **Black box przy złych outputach** - gdy AI generuje słaby output, nie możesz poprawić promptów OPSX to otwiera. Teraz każdy może eksperymentować z instrukcjami, testować granularnie każdy artefakt, customizować workflows i iterować szybko bez rebuild. ```text Legacy workflow: OPSX: ┌────────────────────────┐ ┌────────────────────────┐ │ Hardcoded in package │ │ schema.yaml │◄── You edit this │ (can't change) │ │ templates/*.md │◄── Or this │ ↓ │ │ ↓ │ │ Wait for new release │ │ Instant effect │ │ ↓ │ │ ↓ │ │ Hope it's better │ │ Test it yourself │ └────────────────────────┘ └────────────────────────┘ ``` **Kluczowa różnica:** OPSX to **actions, nie phases**. Rób co potrzebujesz, kiedy potrzebujesz. ## Komendy OPSX - przegląd OPSX daje Ci zestaw komend do różnych momentów pracy. Nie ma przymusu używania ich w określonej kolejności - to toolkit, nie checklist. | Komenda | Co robi | |---------|---------| | `/opsx:explore` | Myślenie, badanie problemu, porównywanie opcji | | `/opsx:new` | Start nowej zmiany | | `/opsx:continue` | Tworzenie kolejnego artefaktu (na podstawie zależności) | | `/opsx:ff` | Fast-forward - wszystkie planning artifacts naraz | | `/opsx:apply` | Implementacja tasks | | `/opsx:sync` | Synchronizacja delta specs do main | | `/opsx:archive` | Archiwizacja po zakończeniu | Typowy flow wygląda tak: ```text # Eksploracja pomysłu /opsx:explore # Start nowej zmiany /opsx:new # Iteracyjne tworzenie artefaktów /opsx:continue # powtórz aż wszystko gotowe # Implementacja /opsx:apply ``` **Pro tip:** Użyj `/opsx:ff` gdy masz jasny obraz tego co chcesz zbudować. `/opsx:continue` jest lepszy przy eksploracji, gdy chcesz iterować po jednym artefakcie naraz. ## Jak działa OPSX - architektura artefaktów Pod spodem OPSX używa **Directed Acyclic Graph (DAG)** do zarządzania artefaktami. Brzmi skomplikowanie, ale koncept jest prosty: artefakty mają zależności które muszą być spełnione zanim można je utworzyć. ```text proposal (root node) │ ┌─────────────┴─────────────┐ │ │ ▼ ▼ specs design (requires: (requires: proposal) proposal) │ │ └─────────────┬─────────────┘ │ ▼ tasks (requires: specs, design) ``` Każdy artefakt może być w jednym z trzech stanów: ```text BLOCKED ────────────────► READY ────────────────► DONE │ │ │ Missing All deps File exists dependencies are DONE on filesystem ``` **Kluczowe koncepty:** - **Dependencies są enablers, nie gates** - pokazują co jest możliwe, nie co wymagane jako następne - **Filesystem jako state** - istnienie pliku na dysku = artefakt DONE - **Topological ordering** - system wie co tworzyć dalej dzięki sortowaniu topologicznemu grafu To oznacza że `/opsx:continue` zawsze wie który artefakt jest następny do utworzenia. Nie musisz pamiętać co już istnieje. ## Customizacja workflow - schematy i konfiguracja OPSX pozwala definiować własne workflows przez **schematy**. Domyślny schemat `spec-driven` wygląda tak: proposal → specs → design → tasks. Ale możesz stworzyć własny. **Przykład własnego schematu:** ```yaml name: research-first artifacts: - id: research generates: research.md requires: [] - id: proposal generates: proposal.md requires: [research] - id: tasks generates: tasks.md requires: [proposal] ``` Ten schemat dodaje `research` przed `proposal` - przydatne gdy pracujesz nad czymś co wymaga zbadania przed commitment do konkretnego rozwiązania. **Konfiguracja projektu** pozwala wstrzyknąć context do wszystkich artefaktów: ```yaml # openspec/config.yaml schema: spec-driven context: | Tech stack: TypeScript, React, Node.js API conventions: RESTful, JSON responses Testing: Vitest for unit tests, Playwright for e2e rules: proposal: - Include rollback plan - Identify affected teams specs: - Use Given/When/Then format ``` **Context injection** sprawia że AI zna konwencje twojego projektu. Zamiast powtarzać "używamy TypeScript, REST API, Vitest" w każdym promcie, definiujesz to raz w configu. ## Kiedy aktualizować istniejącą zmianę vs zacząć nową Możesz zawsze edytować proposal lub specs przed implementacją. Ale kiedy refinement staje się "to już inna praca"? **Proposal definiuje trzy rzeczy:** 1. **Intent** - Jaki problem rozwiązujesz? 2. **Scope** - Co jest in/out of bounds? 3. **Approach** - Jak zamierzasz to rozwiązać? ```text ┌─────────────────────────────────────┐ │ Czy to ta sama praca? │ └──────────────┬──────────────────────┘ │ ┌──────────────────┼──────────────────┐ │ │ │ ▼ ▼ ▼ Same intent? >50% overlap? Można zamknąć Same problem? Same scope? oryginalną zmianę? │ │ │ ┌────────┴────────┐ ┌──────┴──────┐ ┌───────┴───────┐ │ │ │ │ │ │ TAK NIE TAK NIE NIE TAK │ │ │ │ │ │ ▼ ▼ ▼ ▼ ▼ ▼ UPDATE NOWA UPDATE NOWA UPDATE NOWA ``` | Test | Aktualizuj | Nowa zmiana | |------|------------|-------------| | **Tożsamość** | "To samo, dopracowane" | "Inna praca" | | **Overlap scope** | >50% pokrycia | <50% pokrycia | | **Zamknięcie** | Nie można bez zmian | Można zamknąć, nowa stoi samodzielnie | > **Update zachowuje kontekst. Nowa zmiana daje klarowność.** Pomyśl o tym jak o git branches - commituj dopóki pracujesz nad tym samym feature, zacznij nowy branch gdy to genuinely nowa praca. ## Jak zacząć z OPSX Setup jest prosty: ```bash # Instalacja openspec init # Sprawdzenie dostępnych schematów openspec schemas # Status aktywnych zmian openspec status ``` `openspec init` tworzy skills w `.claude/skills/` które AI coding assistants automatycznie wykrywają. **Typowy workflow od zera:** ```text 1. /opsx:explore → przemyśl pomysł 2. /opsx:new → zacznij zmianę 3. /opsx:continue → stwórz proposal 4. /opsx:continue → stwórz specs 5. /opsx:continue → stwórz design 6. /opsx:continue → stwórz tasks 7. /opsx:apply → implementuj 8. /opsx:archive → zakończ ``` **Tips dla startu:** - Użyj `/opsx:explore` przed commitment do zmiany - przemyśl opcje - `/opsx:ff` gdy wiesz co chcesz, `/opsx:continue` przy eksploracji - Podczas `/opsx:apply` - jeśli coś nie tak, edytuj artefakt i kontynuuj (brak phase gates!) ## Kluczowe wnioski 1. **OPSX to actions, nie phases** - rób co potrzebujesz, kiedy potrzebujesz 2. **Artefakty tworzą graf zależności** - system wie co jest ready do utworzenia 3. **Iteracja jest naturalna** - edytuj specs podczas implementacji, to nie bug to feature 4. **Schematy są customizable** - zdefiniuj własny workflow dopasowany do twojego procesu 5. **Context injection** - AI zna konwencje twojego projektu bez powtarzania w każdym promcie

Chcesz wdrożyć OPSX w swoim zespole?

Pomagam zespołom przejść od chaotycznego promptowania do strukturyzowanego workflow z AI. Napisz, a ustalimy czy OPSX pasuje do waszego procesu.

Umów bezpłatną konsultację
## Przydatne zasoby - [OpenSpec GitHub](https://github.com/Fission-AI/openspec) - oficjalne repo projektu - [OpenSpec Discord](https://discord.gg/YctCnvvshC) - community i feedback - [5 technik pracy z Claude Code](/blog/5-technik-pracy-z-claude-code) - powiązany artykuł o AI coding - [Second Brain z Obsidian i Claude Code](/blog/second-brain-obsidian-claude-code-skills) - jak zarządzać kontekstem i skills ## FAQ
### Czy OPSX działa tylko z Claude Code czy też z Cursor i innymi AI coding assistants? OPSX generuje skills do `.claude/skills/` które są cross-editor compatible. Działa z Claude Code, Cursor, Windsurf i innymi asystentami wspierającymi format skills. Kluczowe jest że `openspec init` tworzy odpowiednie pliki dla każdego edytora automatycznie.
### Jaka jest różnica między komendą /opsx:continue a /opsx:ff i kiedy używać której? `/opsx:continue` tworzy jeden artefakt na raz, co jest idealne przy eksploracji gdy chcesz iterować i weryfikować każdy krok. `/opsx:ff` (fast-forward) tworzy wszystkie planning artifacts naraz. Użyj `ff` gdy masz jasny obraz tego co budujesz, `continue` gdy chcesz iterować i przemyśleć każdy artefakt osobno.
### Czy mogę stworzyć własny schemat OPSX dostosowany do workflow mojego zespołu? Tak, użyj `openspec schema init my-workflow` do stworzenia nowego schematu od zera lub `openspec schema fork spec-driven my-workflow` żeby zacząć od istniejącego. Schematy to pliki YAML w `openspec/schemas/` gdzie definiujesz artefakty, ich output i zależności między nimi.
### Jak OPSX radzi sobie z sytuacją gdy podczas implementacji okazuje się że design jest błędny? To jest core feature OPSX - po prostu edytujesz `design.md` bezpośrednio i kontynuujesz. `/opsx:apply` kontynuuje od miejsca gdzie skończyłeś. Brak "phase gates" oznacza że możesz wrócić do dowolnego artefaktu kiedy chcesz bez restartowania całego procesu.
### Czy potrzebuję zainstalowanego CLI openspec żeby używać komend OPSX w edytorze? Tak, skills (`/opsx:*`) wywołują CLI `openspec` w tle. Komendy slash to interface dla użytkownika, CLI to engine który wykonuje faktyczną pracę. Instalacja przez `npm install -g openspec` lub zgodnie z instrukcjami w repo projektu.
--- # Remotion + AI: Jak tworzyć profesjonalne wideo za pomocą kodu i Claude Source: https://pawel.lipowczan.pl/blog/remotion-explainer-videos-ai Published: 2026-01-31 # Remotion + AI: Jak tworzyć profesjonalne wideo za pomocą kodu i Claude Profesjonalne wideo za tysiące złotych? A może w kilka minut i za darmo? Z własnego doświadczenia wiem, że produkcja wideo marketingowego to jedna z najbardziej frustrujących części prowadzenia biznesu. Albo płacisz studiu produkcyjnemu, albo spędzasz tygodnie ucząc się After Effects. Albo po prostu rezygnujesz. Ja wybrałem inną drogę. Stworzyłem **45-sekundowy explainer video** dla mojej strony konsultingowej w kilka minut. Bez Adobe, bez Canvy, bez wiedzy o animacji. Tylko React, **Remotion** i **Claude Code**. Efekt możesz zobaczyć na [konsultacje.lipowczan.pl/explainer](https://konsultacje.lipowczan.pl/explainer). Tradycyjnie tworzenie explainera wygląda tak: wymyślasz koncepcję, piszesz scenariusz, tworzysz storyboard, projektujesz grafiki, animujesz każdy element frame-by-frame, eksportujesz, poprawiasz, eksportujesz znowu. Minimum kilka dni pracy. Realnie - tygodnie. Moja metoda? Opisuję co chcę zobaczyć, AI generuje kod, renderuję wideo. Całość w czasie lunchu. W tym artykule pokażę Ci kompletny workflow - od instalacji przez pisanie promptów po eksport gotowego wideo. Jeśli prowadzisz biznes i potrzebujesz wideo marketingowego, ale nie masz budżetu na studio produkcyjne ani czasu na naukę narzędzi - ten przewodnik jest dla Ciebie. ## Co to jest Remotion? **Remotion** to biblioteka do tworzenia wideo w React. Zamiast przeciągać elementy na timeline jak w Premiere, opisujesz je kodem. Zamiast animować ręcznie, definiujesz reguły - a Remotion generuje płynne przejścia. Brzmi technicznie? W praktyce to prostsze niż tradycyjna edycja wideo. ```text Tradycyjne podejście: Pomysł → Storyboard → Nagranie → Edycja → Eksport (dni/tygodnie) Remotion + AI: Pomysł → Prompt → Kod → Renderowanie (minuty) ``` Różnica jest fundamentalna. W tradycyjnych narzędziach mówisz **jak** animować każdą klatkę. W Remotion opisujesz **co** chcesz zobaczyć - narzędzie samo oblicza interpolację, timing i przejścia. Co więcej, wszystko jest **programowalne**. Chcesz zmienić kolor brandowy? Jedna zmienna. Chcesz wygenerować 50 wersji wideo z różnymi nazwami produktów? Loop w JavaScript. Chcesz dodać dynamiczne dane z API? Import i render. To idealne narzędzie dla podejścia **vibe coding**, które opisywałem [w osobnym artykule](/blog/vibe-coding-przewodnik). Nie musisz być ekspertem od React - wystarczy opisać co chcesz, a AI wygeneruje działający kod. Remotion jest open-source i darmowy do użytku osobistego. Płatna licencja wymagana tylko dla firm z przychodami powyżej $1M rocznie - czyli prawdopodobnie nie musisz się tym przejmować. ## Zobacz efekt: Przykładowy explainer Zanim przejdziemy do instalacji, zobacz co możesz stworzyć. To wideo wygenerowałem w kilka minut używając Claude Code i Remotion: ### Prompt użyty do stworzenia Oto dokładny prompt, którym opisałem wideo: ```markdown Stwórz 45-sekundowe explainer video o tworzeniu wideo z Remotion i AI. **Sceny:** 1. (0-8s) **Problem** - Tekst "Profesjonalne wideo?" z efektem glitch - Pod spodem animowane ikony: zegar + znak dolara + znak zapytania - Fade out do ciemności 2. (8-16s) **Tradycyjne podejście (przekreślone)** - Pozioma timeline z ikonami: Pomysł → Storyboard → Edycja → Export - Czerwona linia przekreśla całość - Tekst "Tygodnie pracy" zanika 3. (16-28s) **Nowe podejście - hero moment** - Duży napis "Remotion + AI" z gradient glow (#00ff9d → #00b8ff) - Trzy animowane karty wlatują od dołu (staggered, 0.3s delay): - "React" z ikoną - "Remotion" z ikoną wideo - "Claude Code" z ikoną AI - Subtelny particle effect w tle 4. (28-38s) **Workflow demo** - Symulacja terminala z kodem (font: Fira Code) - Typing animation: `npx remotion render...` - Progress bar 0% → 100% - Tekst "45 sekund → Gotowe wideo" z pulsującym glow 5. (38-45s) **CTA** - Logo/tekst "pawel.lipowczan.pl" z gradient underline - Tekst "Więcej na blogu" fade in - Subtelna animacja network mesh w tle **Styl wizualny:** - Dark mode, tło: #0a0e1a - Primary accent: #00ff9d (gradient do #00b8ff) - Glassmorphism na kartach (backdrop-blur, border-white/10) - Typography: Inter (headers), Fira Code (kod) - Animacje: smooth easing, spring physics dla wejść elementów **Format:** 16:9 (1920x1080) **Tempo:** Dynamiczne, tech-forward **Muzyka:** Brak (opcjonalnie ambient synth loop) **Brand colors (hex):** - Primary: #00ff9d - Secondary: #00b8ff - Background: #0a0e1a - Dark surface: #151b2b - Text: #ffffff - Muted text: #9ca3af ``` Widzisz jak szczegółowy jest prompt? Każda scena ma timing, opis elementów i efektów. To klucz do dobrych wyników. ### Ważna uwaga: Używaj skilla explicite Podczas testowania zauważyłem istotny szczegół. Gdy wysłałem prompt bez jawnego wskazania skilla `remotion-best-practices`, Claude zaczął skanować całe repozytorium i przygotowywać zbędny plan. Dopiero gdy explicite wywołałem skill (`/remotion-best-practices`), Claude od razu zaczął poprawnie budować projekt z animacjami. **Wniosek:** Jeśli masz zainstalowany skill Remotion, zawsze uruchamiaj go świadomie zamiast liczyć na automatyczne wykrycie kontekstu. ### Komendy do pracy z Remotion Po wygenerowaniu kodu masz trzy opcje: | Komenda | Co robi | | ------------------- | -------------------------------------------------------- | | `npm run dev` | Uruchamia Remotion Studio - podgląd na żywo z hot reload | | `npm run build` | Renderuje wideo do MP4 (domyślnie out/video.mp4) | | `npm run build:gif` | Renderuje animowany GIF (mniejszy plik, niższa jakość) | Remotion Studio to najlepszy sposób na iterację - widzisz zmiany natychmiast i możesz scrubować timeline. ## Instalacja i pierwsze kroki Zanim zaczniesz, potrzebujesz **Claude Code** (CLI Anthropica) lub **Claude Desktop z MCP**. Jeśli jeszcze nie masz skonfigurowanego środowiska, zajrzyj do mojego artykułu o [5 technikach pracy z Claude Code](/blog/5-technik-pracy-z-claude-code). Instalacja składa się z dwóch kroków. Najpierw tworzysz projekt Remotion: ```bash npx create-video@latest ``` To standardowa komenda Remotion, która tworzy nowy projekt z całą strukturą plików - komponenty React, konfigurację i przykładowe wideo. Następnie instalujesz skill dla Claude Code: ```bash npx skills add remotion-dev/skills ``` Co się dzieje podczas instalacji? **Skills** to rozszerzenia Claude Code, które dodają specjalistyczną wiedzę. Skill `remotion-best-practices` uczy AI najlepszych praktyk tworzenia animacji w Remotion - struktury komponentów, timing, interpolacji i eksportu. Jak to działa w praktyce? Po prostu opisujesz wideo, które chcesz stworzyć. Claude automatycznie rozpoznaje kontekst Remotion i wykorzystuje wiedzę ze skilla do wygenerowania kodu. Nie ma specjalnych komend do zapamiętania - piszesz prompt naturalnym językiem. Przykładowy workflow: ```text 1. Otwórz projekt Remotion w Claude Code 2. Opisz wideo: "Stwórz 30-sekundową animację logo z efektem glow" 3. Claude generuje komponenty React z animacjami 4. Uruchom podgląd: npx remotion studio 5. Wyrenderuj: npx remotion render src/index.ts MyVideo out/video.mp4 ``` **Ważne:** Nie musisz znać React na zaawansowanym poziomie. Wystarczy podstawowe zrozumienie komponentów i props. Resztę zrobi AI. W praktyce moje prompty nie zawierają ani linijki kodu - opisuję tylko efekt wizualny. ## Prompt engineering dla wideo Tu zaczyna się właściwa praca. Jakość wideo zależy bezpośrednio od jakości promptu. AI nie czyta w myślach - potrzebuje konkretnych instrukcji. ### Anatomia dobrego promptu Każdy dobry prompt do wideo zawiera cztery elementy: 1. **Podział na sceny z timingiem** - AI musi wiedzieć co dzieje się w której sekundzie 2. **Styl wizualny** - kolory, nastrój, estetyka 3. **Brand assets** - konkretne wartości hex, nie "zielony" 4. **Format i przeznaczenie** - 16:9 dla YouTube, 9:16 dla Stories Przykład strukturalnego promptu: ```markdown Stwórz 45-sekundowe video explainer dla strony konsultingowej. **Sceny:** 1. (0-10s) Logo animacja z gradientem #00ff9d → #00cc7d 2. (10-25s) 3 główne usługi jako animowane karty 3. (25-40s) Testimonial lub statystyka 4. (40-45s) CTA z kontaktowym przyciskiem **Styl:** Profesjonalny, minimalistyczny, dark mode **Format:** 16:9 (YouTube/website) **Tempo:** Spokojne, business-oriented ``` Widzisz różnicę między "zrób ładne wideo o mojej firmie" a tym promptem? Pierwszy da losowe wyniki. Drugi da dokładnie to, czego potrzebujesz. ### Szablon promptu Pawła Ten prompt użyłem do stworzenia explainera dla mojej strony. Możesz go skopiować i dostosować: ```markdown For the current project create a 45-second explainer video based on the content. The video should: - Extract and use the brand colors from the site - Highlight the main services/products - Include key selling points - Use professional, elegant animations - End with a clear call-to-action Aspect ratio: 16:9 Style: PROFESSIONAL ``` Ten prompt działa szczególnie dobrze, gdy masz już stronę lub design system. Claude Code analizuje kontekst projektu i wyciąga kolory, typografię i styl z istniejących plików. Kluczowe frazy, które pomagają: - **"Extract from site"** - AI przeszuka projekt i użyje istniejących wartości - **"Professional, elegant"** - vibe descriptors są skuteczniejsze niż szczegółowe instrukcje - **"Clear call-to-action"** - AI wie, że końcówka ma skłaniać do działania ## Przypadki użycia Remotion + AI to nie tylko explainery. Oto praktyczne zastosowania, które przetestowałem lub widziałem w akcji. ### Dla przedsiębiorców i konsultantów **Explainer videos** - krótkie wideo wyjaśniające usługę. Idealne na landing page, gdzie 30-60 sekund animacji może zastąpić paragraf tekstu. **Product demos** - pokaż jak działa Twój produkt bez nagrywania ekranu. Animowane mockupy wyglądają profesjonalniej i nie starzeją się przy każdej aktualizacji UI. **Case study visualizations** - zamiast nudnej prezentacji PowerPoint, animuj statystyki i wykresy. "Revenue increased 340%" ma większy impact jako animacja niż jako bullet point. ### Dla twórców treści **Animated title sequences** - intro do YouTube videos. Zamiast kupować template na Envato, wygenerujesz unikalny opener pasujący do Twojego brandu. **Lower thirds** - napisy z imieniem i tytułem. Raz wygenerowane, możesz używać wielokrotnie z różnymi danymi. **Social media content** - Stories i Reels w formacie 9:16. Remotion renderuje w dowolnym aspect ratio. ### Dla SaaS i produktów **Feature announcements** - nowa funkcja zasługuje na więcej niż post na blogu. 15-sekundowy teaser wygląda profesjonalnie i generuje zaangażowanie. ```markdown Stwórz 15-sekundowy teaser nowej funkcji dla LinkedIn. Scena 1: Nazwa funkcji jako duży napis z glow effect Scena 2: 3 bullet points z korzyściami (staggered animation) Scena 3: "Już dostępne" z logo produktu Format: 1:1 (LinkedIn feed) Kolory: Paleta produktu (#1a1a2e, #16213e, #0f3460, #e94560) ``` **Onboarding videos** - krótkie animacje wyjaśniające kolejne kroki. Łatwiejsze w aktualizacji niż nagrania z lektorem. **Release notes as video** - changelog w formie animowanej listy. Użytkownicy chętniej obejrzą 30-sekundowe wideo niż przeczytają listę zmian. ## Wskazówki dla lepszych wyników Po kilkunastu projektach wideo z AI zebrałem zestaw praktyk, które konsekwentnie poprawiają wyniki. 1. **Zawsze podawaj timing** - AI domyślnie robi za długie sceny. Bez konkretnych sekund dostaniesz 3-minutowe wideo zamiast 45-sekundowego. 2. **Opisuj brand assets konkretnie** - kolory hex (#00ff9d), nie "zielony". AI interpretuje "niebieski" na 50 różnych sposobów. 3. **Używaj "vibe" opisów** - "Apple-like minimalism" jest bardziej skuteczne niż szczegółowe instrukcje o kerning i leading. AI zna estetykę topowych brandów. 4. **Iteruj małymi krokami** - jedno wideo, jedna zmiana. Nie próbuj naprawiać wszystkiego jednym promptem. "Make the logo animation faster" > "Fix the video". 5. **Testuj różne formaty** - 16:9 ≠ 9:16 w strukturze treści. Pionowe wideo wymaga innego układu elementów, nie tylko przycięcia. 6. **Eksportuj w odpowiedniej jakości** - 1080p dla web, 4K dla prezentacji. Większa rozdzielczość = dłuższy render, ale YouTube i tak kompresuje. Jeśli chcesz więcej o AI-driven animations, sprawdź mój artykuł o [tworzeniu animacji w stylu Apple](/blog/animacje-apple-ai-cursor). Opisuję tam podobny workflow z Google Flow i Cursor. ## Czego Remotion nie zrobi Bądźmy realistami - Remotion + AI nie zastępuje studia produkcyjnego w każdym scenariuszu. **Ograniczenia:** - **Complex motion graphics** - zaawansowane efekty 3D, particle systems, morphing są poza zasięgiem. To nie After Effects. - **Real footage editing** - Remotion nie jest edytorem wideo. Nie przytnie Ci nagrania z kamery. - **Voice generation** - biblioteka nie generuje głosu. Potrzebujesz zewnętrznego TTS (ElevenLabs, Murf, etc.). - **Photo-realistic graphics** - generujesz animacje, nie filmy. Bez kamer, aktorów, planu zdjęciowego. **Obejścia:** - **Łącz z zewnętrznymi narzędziami** - ElevenLabs generuje lektor, Remotion animację. Scalasz w Premiere lub DaVinci. - **Używaj jako warstwa animacji** - Remotion świetnie nakłada animowane elementy na istniejące wideo. - **Hybrid approach** - proste sceny w Remotion, złożone w tradycyjnych narzędziach. Mieszaj i łącz. Remotion jest idealny dla: explainers, motion graphics, data visualization, UI animations, product teasers. Nie jest idealny dla: filmów dokumentalnych, vlogów, content z real footage. ## Podsumowanie Stworzenie profesjonalnego wideo marketingowego nie musi kosztować tysięcy złotych ani zabierać tygodni. Z Remotion i Claude Code możesz mieć gotowy explainer w czasie lunchu. **Kluczowe wnioski:** 1. **Remotion + AI = wideo w minutach zamiast dni** - deklaratywne podejście eliminuje ręczną animację 2. **Nie musisz być animatorem ani programistą** - AI generuje kod, Ty opisujesz efekt 3. **Dobre prompty = dobre wyniki** - struktura scen, timing, kolory hex, vibe descriptors 4. **Zaczynaj od prostych projektów** - logo animation, title card, potem explainer 5. **Iteruj** - AI nie musi trafić za pierwszym razem, ale trafi za trzecim 6. **Format ma znaczenie** - 16:9 vs 9:16 vs 1:1 wymaga różnego podejścia do layoutu Demokratyzacja produkcji wideo to jeden z najciekawszych trendów AI. Jeszcze rok temu byłem przekonany, że wideo to domena specjalistów z Adobe. Dziś wiem, że każdy z dostępem do Claude Code może stworzyć coś, co wygląda profesjonalnie. Zacznij od czegoś małego - animowane logo, 10-sekundowy teaser. Poczuj narzędzie. Potem rozbudowuj.

Chcesz tworzyć profesjonalne wideo dla swojego biznesu?

Pomogę Ci wdrożyć AI w procesie tworzenia treści marketingowych - od strategii przez implementację po automatyzację produkcji.

Umów bezpłatną konsultację
## Zasoby **Oficjalne źródła:** - [Remotion Official Docs](https://remotion.dev/docs) - Oficjalna dokumentacja biblioteki - [Remotion Skills Repository](https://github.com/remotion-dev/skills) - Instalacja i przykłady Skills dla Claude Code - [Mój explainer video](https://konsultacje.lipowczan.pl/explainer) - Efekt końcowy workflow z tego artykułu **Powiązane artykuły:** - [Animacje w stylu Apple z AI](/blog/animacje-apple-ai-cursor) - Podobny workflow z Google Flow i Cursor - [Vibe Coding przewodnik](/blog/vibe-coding-przewodnik) - Filozofia kodowania z AI, która leży u podstaw tego podejścia ## FAQ
### Czy muszę znać programowanie żeby używać Remotion z AI? Nie, Claude Code generuje kod za Ciebie. Wystarczy umiejętność pisania dobrych promptów i podstawowe zrozumienie efektu, który chcesz osiągnąć. React działa "pod maską" - Ty opisujesz sceny, timing i styl, AI tłumaczy to na działający kod.
### Ile kosztuje tworzenie wideo w Remotion i czy potrzebuję płatnej licencji? Remotion jest open-source i darmowy do użytku osobistego oraz komercyjnego dla firm z przychodami poniżej $1M rocznie. Płatna licencja wymagana tylko dla większych organizacji. Claude Code wymaga subskrypcji Claude Pro ($20/miesiąc) lub używania przez API z własnym kluczem.
### Czy mogę użyć wygenerowanego wideo komercyjnie i kto jest właścicielem praw? Tak, wideo stworzone w Remotion należy w pełni do Ciebie. Możesz używać go komercyjnie bez ograniczeń. Upewnij się tylko, że używane assets (czcionki, ikony, muzyka) mają odpowiednie licencje - AI może zasugerować zasoby wymagające dodatkowych praw.
### Jak długo trwa renderowanie 45-sekundowego wideo i od czego to zależy? Typowo 2-5 minut dla prostych animacji na nowoczesnym laptopie z 8GB+ RAM. Czas zależy od złożoności efektów, rozdzielczości (4K dłużej niż 1080p) i liczby elementów. Dla dużych projektów można renderować w chmurze przez Remotion Lambda - szybciej, ale płatne.
### Czy Remotion może wygenerować wideo z lektorem i jak dodać głos do animacji? Remotion nie generuje głosu - to biblioteka do animacji, nie TTS. Możesz połączyć go z ElevenLabs, Murf lub innymi usługami text-to-speech, wygenerować plik audio, a następnie zsynchronizować go z animacją w Remotion. AI pomoże Ci napisać kod synchronizacji audio z scenami.
### Czym różni się Remotion od Canva czy CapCut i kiedy wybrać które narzędzie? Canva i CapCut to edytory wizualne z gotowymi szablonami - świetne dla jednorazowej szybkiej edycji bez pisania kodu. Remotion to narzędzie programistyczne dające pełną kontrolę, powtarzalność i automatyzację. Wybierz Remotion gdy potrzebujesz generować wiele wariantów wideo, integrować z danymi z API lub budować spójny system produkcji wideo.
--- # Jak tworzyć animacje w stylu Apple z pomocą AI Source: https://pawel.lipowczan.pl/blog/animacje-apple-ai-cursor Published: 2026-01-26 # Jak tworzyć animacje w stylu Apple z pomocą AI Animacje na stronach Apple to standard, do którego wszyscy dążymy. Problem w tym, że ich stworzenie wymaga miesięcy nauki After Effects, Lottie i zaawansowanego JavaScriptu. Albo wymagało - do teraz. Ja też nie jestem animatorem. Nie spędziłem lat ucząc się motion design. Ale dzięki narzędziom AI mogę tworzyć **scroll-animacje**, które wyglądają jak prosto ze stron produktowych Apple. Inspirację do tego workflow znalazłem w społeczności [AI Launchpad](https://www.skool.com/signup?ref=965de823f433460284e52128d441d65e) - grupie ponad 12 000 osób eksperymentujących z AI w praktyce. To tam zobaczyłem, jak połączyć Google Whisk, Flow i Cursor w spójny pipeline produkcyjny. W tym artykule pokażę Ci kompletny workflow - od researchu przez generowanie obrazów, tworzenie animacji, aż po deployment działającej strony. Jako przykład użyję **hero section dla strony usług AI** - coś, co sam mógłbym wykorzystać. Jeśli kiedykolwiek chciałeś stworzyć stronę produktową z płynnymi animacjami, ale brakowało Ci umiejętności technicznych - ten przewodnik jest dla Ciebie. ## Dlaczego animacje na stronach są takie trudne? Tradycyjne podejście do tworzenia scroll-animacji to prawdziwy maraton: 1. After Effects do stworzenia animacji 2. Eksport do formatu Lottie 3. Integracja z kodem JavaScript 4. Optymalizacja wydajności 5. Synchronizacja z pozycją scrolla Każdy z tych kroków ma swoją krzywą uczenia się. A do tego dochodzą problemy: - **Scroll-triggered animations** wymagają precyzyjnego timing'u - Animacje muszą być responsywne na różnych urządzeniach - Optymalizacja **frame-by-frame** pod kątem wydajności - Synchronizacja z pozycją scrollowania bez lag'ów Efekt? Większość stron wygląda generycznie. Bo tworzenie czegoś naprawdę wow wymaga połączenia umiejętności designera, animatora i developera. AI zmienia tę grę kompletnie. ## Kompletny workflow AI dla animacji ### Krok 1: Research - zrozum co działa Zanim zaczniesz promptować AI, musisz wiedzieć czego chcesz. Research to fundament. Przejrzyj strony produktowe topowych brandów. Zwróć uwagę na: - Paletę kolorów i kontrast - Strukturę sekcji i pozycjonowanie tekstu - Jak animacje reagują na scroll - Timing i pacing przejść Zapisuj screenshoty jako reference. AI daje generyczne wyniki bez konkretnych inspiracji. Z referencjami - tworzy coś unikalnego. ### Krok 2: Generowanie obrazów (Google Whisk) Teraz czas stworzyć statyczne klatki, które później ożywisz. Używam do tego [Google Whisk](https://labs.google/fx/tools/whisk) - narzędzia, które pozwala generować obrazy z referencjami stylu. Prompt powinien być konkretny i opisowy. Dla wizualizacji sieci neuronowej: ```text Abstract neural network visualization for AI consulting, dark gradient #0a0e1a to #151b2b, interconnected nodes in bright green #00ff9d, synaptic connections in cyan #00b8ff, floating hexagonal elements with glassmorphism, futuristic professional aesthetic, no text ``` Kluczowe elementy dobrego promptu: - Określ konkretne kolory (hex codes) - Opisz styl (futuristic, professional, glassmorphism) - Dodaj negatywne instrukcje (no text, no letters) - Wskaż kompozycję i nastrój **Opcja A: Dwie klatki + Google Flow** Wygeneruj dwie klatki: początkową i końcową. Na przykład: statyczna sieć neuronowa → aktywna sieć z pulsującymi połączeniami. Potem użyj Google Flow (krok 3) do stworzenia przejścia. **Opcja B: Obraz → wideo w Whisk (szybsza)** Whisk pozwala też wygenerować wideo bezpośrednio z obrazu. Stwórz jeden obraz, a następnie kliknij ikonę wideo i opisz ruch. Whisk wygeneruje animację na podstawie statycznej klatki - pomijasz wtedy krok 3. **Kiedy wybrać którą opcję?** - Wybierz **opcję A**, gdy chcesz mieć **większą kontrolę nad przejściem** (np. wyraźny stan początkowy i końcowy, konkretny „storytelling” między klatkami) i zależy Ci na bardziej dopracowanej, filmowej animacji. - Wybierz **opcję B**, gdy potrzebujesz **szybkiego rezultatu z jednego obrazu**, chcesz tylko „ożywić” statyczną grafikę lub szybko przetestować pomysł na ruch bez przygotowywania dwóch klatek. ### Krok 3: Tworzenie animacji (Google Flow) - opcjonalny [Google Flow](https://labs.google/fx/tools/flow) to narzędzie, które tworzy płynne przejścia między klatkami. Upload dwóch obrazów (początkowy + końcowy) i opisz przejście: ```text Neural network awakening, nodes lighting up sequentially, data flowing through connections, synaptic pulses, smooth cinematic transition, professional tech aesthetic ``` AI wygeneruje wideo z płynnym przejściem między klatkami. Magia dzieje się automatycznie - nie musisz animować każdego elementu osobno. ### Krok 4: Konwersja wideo na sekwencję obrazów Tu jest clue całego workflow. Scroll-animacje nie pracują z wideo - pracują z sekwencjami obrazów, które wyświetlają się w reakcji na pozycję scrolla. Narzędzie: [Online Convert](https://image.online-convert.com/convert/mp4-to-jpg) (darmowe) 1. Upload wideo z Google Flow lub Whisk 2. Wybierz liczbę klatek (24-60 fps) 3. Pobierz archiwum z JPEG frames Rezultat: folder z dziesiątkami obrazów, które tworzą płynną animację gdy wyświetlane sekwencyjnie. ## Cursor: AI IDE do budowania stron z animacjami **[Cursor](https://cursor.sh)** to AI-powered IDE, które pozwala budować kompletne strony przez rozmowę z AI. W trybie Composer możesz opisać czego potrzebujesz, a Cursor wygeneruje działający kod. ### Konfiguracja projektu (.cursorrules) Zanim zaczniesz budować, stwórz plik `.cursorrules` w głównym folderze projektu. To zasady, których AI będzie przestrzegać w każdym prompcie: ```text 1. Always use semantic HTML5 elements for accessibility and SEO 2. Maintain consistent spacing using 8px grid system 3. All animations should respect prefers-reduced-motion for accessibility 4. Use CSS variables for colors to enable easy theme switching 5. Use Framer Motion for scroll-triggered animations 6. Image sequence should be loaded progressively for performance ``` Dlaczego to ważne? Jednorazowa konfiguracja, permanentne korzyści. Każdy komponent będzie spójny bez powtarzania tych instrukcji. ### Strukturalny prompt dla fundamentu strony W trybie Composer (Cmd+I / Ctrl+I) opisz strukturę strony: ```text Create an AI consulting landing page with React and Tailwind: - Dark theme (#0a0e1a to #151b2b gradient) - Hero section with neural network animation - Services section highlighting AI automation - Case studies carousel - CTA for consultation booking Style: futuristic, professional, tech-forward Typography: Inter for body, bold geometric headlines Color accents: #00ff9d (green), #00b8ff (cyan) ``` Cursor wygeneruje kompletną strukturę strony z komponentami React. Hero, services, case studies, CTA - wszystko gotowe do customizacji. ### Integracja animacji ze scroll-triggerem Teraz najważniejszy krok. Dodaj wszystkie frames do folderu `public/frames/` i użyj promptu w Cursor: ```text Create a scroll-triggered animation component using Framer Motion. Load image sequence from /frames/ folder (frame-001.webp to frame-060.webp). Animation should play forward as user scrolls down, reverse when scrolling up. Use useScroll and useTransform hooks for smooth interpolation. Full viewport height hero section with centered text overlay. ``` Cursor wygeneruje komponent z Framer Motion, który synchronizuje wyświetlanie klatek z pozycją scrolla - dokładnie jak na stronach Apple. ### Dodatkowe usprawnienia Po podstawowej konfiguracji możesz dopracować szczegóły przez kolejne prompty: - **Preloading**: "Add progressive image preloading for smoother animation" - **Reduced motion**: "Add support for prefers-reduced-motion media query" - **Typography**: "Use variable font with responsive sizing" Każdy prompt dopracowuje stronę. Iteracja to klucz - Cursor pamięta kontekst całego projektu. ## Publikacja strony - Netlify w 5 minut Masz gotową stronę? Czas ją opublikować. ### Build projektu W terminalu Cursor uruchom build: ```bash npm run build ``` Rezultat: folder `dist` z produkcyjnym buildem strony, zoptymalizowanymi obrazami i minifikowanym kodem. ### Deploy na Netlify Netlify oferuje najprostszy deployment dla static sites: 1. Wejdź na [netlify.com](https://netlify.com) 2. Przeciągnij folder `dist` na stronę (drag & drop) 3. Gotowe - strona jest live Alternatywnie połącz z GitHub repo dla automatycznych deployments przy każdym push'u. Custom domain? Netlify obsługuje to natywnie - dodaj DNS records i masz profesjonalny adres. ## Cursor vs ręczne kodowanie - kiedy co wybrać? | Kryterium | Cursor + AI Images | Ręczne kodowanie | | ----------------- | -------------------------------- | ----------------- | | Poziom techniczny | Początkujący/Średni | Zaawansowany | | Kontrola | Wysoka (kod jest Twój) | Pełna | | Customizacja | Przez prompty + edycję | Przez kod | | Czas realizacji | 30-60 minut | Kilka godzin | | Najlepsze dla | Szybkie prototypy, landing pages | Złożone aplikacje | Na mojej stronie portfolio używam Framer Motion pisanego ręcznie, bo potrzebuję pełnej kontroli nad każdą animacją. Ale gdybym budował landing page dla klienta? Cursor przyspiesza pracę znacząco - generuję szkielet, potem dopracowuję szczegóły. Jeśli chcesz dowiedzieć się więcej o podejściu AI do tworzenia UI, sprawdź mój artykuł o [Vibe Coding](/blog/vibe-coding-przewodnik). ## Kluczowe wnioski 1. **Research przed promptowaniem** - bez referencji AI daje generyczne wyniki 2. **Sekwencja obrazów > wideo** - tak działają scroll-animacje w praktyce 3. **.cursorrules to game changer** - jednorazowe ustawienie, permanentne korzyści 4. **Accessibility matters** - `prefers-reduced-motion` to nie opcja, to standard 5. **Deploy jest prosty** - Netlify manual upload to dosłownie drag & drop 6. **Vibe coding to nie magia** - to metodyczne przygotowanie + AI execution

Potrzebujesz strony z animacjami, które robią wrażenie?

Pomagam firmom tworzyć strony internetowe z płynnymi animacjami i profesjonalnym designem - od koncepcji po wdrożenie.

Umów bezpłatną konsultację
## Przydatne zasoby - [AI Launchpad](https://www.skool.com/signup?ref=965de823f433460284e52128d441d65e) - społeczność 12k+ osób eksperymentujących z AI, inspiracja dla tego workflow - [Google Whisk](https://labs.google/fx/tools/whisk) - generowanie obrazów AI z referencjami stylu - [Google Flow](https://labs.google/fx/tools/flow) - tworzenie płynnych przejść między klatkami - [Cursor](https://cursor.sh) - AI-powered IDE do budowania stron i aplikacji - [Online Convert](https://image.online-convert.com/convert/mp4-to-jpg) - konwersja wideo na sekwencję obrazów - [Netlify](https://netlify.com) - darmowy hosting dla static sites - [Framer Motion](https://motion.dev) - biblioteka animacji dla React ## FAQ
### Czy Cursor jest darmowy i jakie ma ograniczenia? Cursor oferuje darmowy tier z limitem zapytań do AI (około 50 wolniejszych zapytań miesięcznie). Do tego workflow wystarczy wersja darmowa. Plan Pro ($20/miesiąc) daje nielimitowany dostęp do szybkich modeli AI i priorytetowe odpowiedzi. Sprawdź aktualny cennik na cursor.sh - modele cenowe AI tools zmieniają się często.
### Ile klatek (frames) potrzebuję do płynnej scroll-animacji na stronie? Dla płynnej animacji potrzebujesz 24-60 klatek na sekundę animacji. W praktyce oznacza to 30-90 obrazów dla 2-3 sekundowej sekwencji. Więcej klatek = płynniejsza animacja, ale większy rozmiar strony. Kompromis: 30-45 frames w formacie WebP daje dobry balans między płynnością a wydajnością.
### Czy mogę użyć własnych zdjęć zamiast obrazów generowanych przez AI? Tak, workflow działa identycznie z własnymi zdjęciami. Przygotuj zdjęcia początkowe i końcowe, użyj Google Flow do wygenerowania przejścia, przekonwertuj na sekwencję obrazów. Własne zdjęcia dają bardziej autentyczny look - szczególnie dla istniejących produktów lub brandingu, gdzie AI nie odtworzy dokładnego wyglądu.
### Jak zoptymalizować scroll-animacje pod kątem wydajności na urządzeniach mobilnych? Trzy kluczowe optymalizacje: konwertuj wszystkie obrazy do formatu WebP (70-80% mniejszy rozmiar), używaj lazy loading dla frames poza viewport, dodaj CSS media query `prefers-reduced-motion: reduce` dla użytkowników z ustawieniami dostępności. Przy generowaniu kodu przez Cursor, dodaj te wymagania do promptu lub .cursorrules.
### Czy potrzebuję umiejętności kodowania, żeby stworzyć stronę z animacjami w tym workflow? Podstawowe umiejętności są przydatne - musisz uruchomić `npm install` i `npm run build` w terminalu. Cursor generuje kod, ale warto rozumieć co robi. Deployment na Netlify to drag & drop. Do prostej strony ze scroll-animacjami wystarczą umiejętności opisane w tym przewodniku - Cursor wyjaśni kod na życzenie.
--- # Second Brain z Obsidian i Claude Code - jak AI zmienia organizację wiedzy Source: https://pawel.lipowczan.pl/blog/second-brain-obsidian-claude-code-skills Published: 2026-01-26 Claude Code to nie tylko narzędzie do kodowania. Brzmi jak clickbait, ale to jedna z najważniejszych rzeczy, które zrozumiałem w ostatnich miesiącach. Odkryłem to dzięki Cole Medin i jego podejściu do używania Claude Code dosłownie do wszystkiego - od zarządzania notatkami po generowanie contentu. Problem, który pewnie znasz: notatki rozrzucone po dziesiątkach narzędzi, historia czatów z AI znika po każdej sesji, a przełączanie między aplikacjami zabija produktywność. Szukałem sposobu na **organizację wiedzy**, który nie wymaga ciągłego ręcznego porządkowania. W tym artykule pokażę Ci jak połączenie **Obsidian**, **Claude Code** i **Skills** tworzy potężny system zarządzania wiedzą. To nie teoria - używam tego setup'u codziennie. Na końcu będziesz mieć wszystko co potrzebne, żeby zbudować własny **second brain** z AI w centrum. ## Czym jest Second Brain i dlaczego go potrzebujesz **Second Brain** to zewnętrzny system do przechowywania i organizacji wiedzy. Koncepcja wywodzi się z **PKM (Personal Knowledge Management)** - podejścia do świadomego zbierania, organizowania i wykorzystywania informacji. Tradycyjny second brain opiera się na trzech funkcjach: - **Capture** - szybkie przechwytywanie myśli, notatek, pomysłów - **Organize** - kategoryzacja i łączenie informacji - **Retrieve** - odnajdywanie wiedzy gdy jej potrzebujesz Problem? Te trzy funkcje wymagają dużo manualnej pracy. Musisz sam decydować gdzie umieścić notatkę, jakie tagi dodać, jak połączyć z innymi dokumentami. AI zmienia zasady gry. Zamiast ręcznie organizować, możesz **poprosić Claude Code o przetworzenie notatek**. Zamiast szukać połączeń samemu, AI analizuje Twój vault i znajduje powiązania. Zamiast tworzyć dokumenty od zera, generujesz je z istniejących notatek. Second brain z AI to nie tylko miejsce do przechowywania wiedzy. To **system, który aktywnie pomaga Ci z tej wiedzy korzystać**. ## Dlaczego Obsidian i Claude Code to idealne połączenie ### Obsidian jako fundament **Obsidian** to edytor notatek oparty na plikach markdown. Kluczowe cechy które czynią go idealnym fundamentem: - **Pliki lokalne** - notatki to zwykłe pliki `.md` na Twoim komputerze - **Offline-capable** - nie potrzebujesz internetu żeby pracować - **Markdown** - format który LLM-y rozumieją najlepiej - **Graph view** - wizualizacja połączeń między notatkami - **Brak vendor lock-in** - pliki są Twoje, możesz je otworzyć w dowolnym edytorze To ostatni punkt jest kluczowy. W przeciwieństwie do Notion czy Evernote, Twoje notatki w Obsidian to **zwykłe pliki tekstowe**. Claude Code może je bezpośrednio czytać, edytować i tworzyć nowe. ### Claude Code jako mózg **Claude Code** to CLI do pracy z Claude. Ale to znacznie więcej niż narzędzie do kodowania. Jego możliwości wykraczają daleko poza pisanie kodu: - **File operations** - czytanie, edycja, tworzenie plików - **Search** - przeszukiwanie zawartości i struktury projektu - **Terminal commands** - uruchamianie skryptów, narzędzi - **Web search** - research bezpośrednio z terminala Tylko **code intelligence** (rozumienie składni, sugestie refactoringu) jest specyficzne dla kodowania. Reszta? To uniwersalne możliwości asystenta AI który ma dostęp do Twojego systemu plików. ### Razem - coś więcej niż suma części Połączenie Obsidian + Claude Code daje coś, czego nie osiągniesz z innymi narzędziami. Notion wymaga MCP servera żeby połączyć się z AI. Obsidian? Claude Code po prostu otwiera folder i czyta pliki. To oznacza: - **Bezpośredni dostęp** do wszystkich notatek bez dodatkowej konfiguracji - **AI przetwarza Twoją bazę wiedzy** - tworzy podsumowania, łączy idee - **Strukturyzowane outputy** z chaotycznych notatek - **Automatyczne połączenia** między dokumentami Przykładowa struktura folderu second brain: ```text obsidian-vault/ ├── 00-inbox/ # Quick captures ├── 01-projects/ # Active projects ├── 02-areas/ # Ongoing responsibilities ├── 03-resources/ # Reference material ├── 04-archive/ # Completed items ├── templates/ # Document templates └── .claude/ └── skills/ # Claude Code skills ``` Ta struktura oparta jest na **PARA method** (Projects, Areas, Resources, Archive). Folder `.claude/skills/` to miejsce gdzie definiujesz capabilities dla Claude Code. ## Skills - jak rozszerzyć możliwości swojego second brain **Skills** to trzeci filar systemu który wszystko spaja. To sposób na dawanie Claude Code wiedzy, procesów i wytycznych specyficznych dla Twojego workflow. ### Co to są Skills Skills to pliki markdown definiujące workflow. Zawierają: - Opis kiedy skill ma być użyty (trigger) - Instrukcje krok po kroku - Referencje do innych plików (templates, style guides) - Przykłady użycia Skill jest **ładowany dynamicznie** gdy Claude Code wykryje że jest potrzebny. Nie musisz go wywoływać ręcznie - po prostu opisujesz co chcesz zrobić. ### Progressive Disclosure - klucz do efektywności To koncepcja która odróżnia skills od innych podejść do rozszerzania AI. Problem z MCP servers: ładują **wszystkie narzędzia z góry**. Jeśli masz 50 narzędzi, wszystkie ich opisy zajmują miejsce w context window. To **context bloat** - marnujesz cenne miejsce na rzeczy których nie używasz. Skills działają inaczej dzięki **progressive disclosure**: 1. **Description** - krótki opis zawsze widoczny (jedna linijka) 2. **SKILL.md** - pełne instrukcje ładowane tylko gdy potrzebne 3. **Reference files** - dodatkowe zasoby ładowane dla konkretnych operacji To oznacza że możesz mieć **setki skills** bez przytłaczania context window. Agent specjalizuje się per sesja - ładuje tylko to co potrzebuje. ### Struktura skill'a ```text .claude/skills/ └── document-generator/ ├── SKILL.md # Main instructions ├── assets/ │ └── templates/ # Document templates └── references/ └── style-guide.md # Style guidelines ``` Przykładowy plik SKILL.md: ```yaml --- name: document-generator description: Generate structured documents from notes. Use when user asks to create summaries, outlines, or formatted documents from their knowledge base. --- # Document Generator ## Triggers - "create a summary of..." - "generate a document from..." - "summarize my notes on..." ## Workflow 1. Read source notes from specified location 2. Analyze key concepts and structure 3. Apply template from `assets/templates/` 4. Generate document following style guide 5. Save to specified location ## Templates Available - `meeting-notes.md` - Meeting summary template - `project-brief.md` - Project overview template - `article-draft.md` - Blog article draft template ``` ## Praktyczne przykłady Skills dla second brain ### Research Engine Skill do zbierania informacji z internetu i zapisywania strukturyzowanych notatek. **Co robi:** - Przeszukuje web na zadany temat - Zbiera kluczowe informacje i źródła - Tworzy notatkę w formacie markdown - Zapisuje do odpowiedniego folderu w vault **Przykład użycia:** "Research the latest developments in AI agents and save notes to 03-resources/ai-agents/" ### Document Generator Transformuje surowe notatki w dopracowane dokumenty. **Co robi:** - Czyta wskazane notatki źródłowe - Analizuje strukturę i kluczowe koncepty - Aplikuje template i style guide - Generuje gotowy dokument **Przykład:** Notatki ze spotkania → dokument z action items i podsumowaniem. ### Daily Review Automatyzuje przegląd dnia. **Co robi:** - Skanuje notatki z dzisiaj - Tworzy podsumowanie dnia - Identyfikuje połączenia z istniejącą wiedzą - Sugeruje next actions ### Content Creator Generuje content na podstawie notatek z vault'a. **Co robi:** - Analizuje notatki na dany temat - Generuje pomysły na content - Tworzy drafty (artykuły, posty, scripts) - Utrzymuje spójny głos i styl Przykład jak wygląda trigger research skill: ```text User: "Research the latest developments in AI agents and save notes to 03-resources/ai-agents/" Claude Code: 1. Loads research-engine skill 2. Searches web for recent AI agent news 3. Creates structured note with sources 4. Saves to specified folder 5. Links to related existing notes ``` ## Łączenie second brain z innymi narzędziami przez MCP **MCP (Model Context Protocol)** to sposób na łączenie Claude Code z zewnętrznymi serwisami. Możesz połączyć się z Gmail, kalendarzem, task managerami. Problem? MCP servers ładują wszystkie narzędzia z góry - to ten sam problem context bloat który rozwiązują skills. ### Dwa podejścia do integracji **1. MCP servers** - bezpośrednie połączenie z usługami - Gmail/Outlook - przez biblioteki lub serwery MCP - Google Calendar - MCP server - ClickUp - MCP server lub API **2. Skills ze skryptami** (moje preferowane podejście) - Piszesz własne skrypty w Python/Node.js - Dodajesz je do skills jako narzędzia - Pełna kontrola nad logiką - Wszystko w jednym miejscu (w vault) ### Dlaczego wolę skrypty w skills Make.com pozwala wystawić scenariusze jako narzędzia MCP. Ale to dodatkowa warstwa. Skrypty w skills są szybsze, prostsze do debugowania, i wszystko mam w jednym miejscu. Przykład skill ze skryptem do ClickUp: ```text .claude/skills/ └── clickup-tasks/ ├── SKILL.md └── scripts/ └── get_tasks.py # Skrypt pobierający taski ``` Plik SKILL.md z użyciem skryptu: ```yaml --- name: clickup-tasks description: Pobierz i zarządzaj taskami z ClickUp. Użyj gdy user pyta o zadania, deadline'y lub status projektów. --- # ClickUp Tasks ## Workflow 1. Uruchom `scripts/get_tasks.py` z odpowiednimi parametrami 2. Przetwórz wyniki i zapisz do notatki 3. Opcjonalnie: połącz z istniejącymi notatkami projektu ## Dostępne operacje - Pobierz taski z listy/folderu - Filtruj po statusie, assignee, deadline - Sync tasków do notatek Obsidian ``` ## Mój osobisty workflow z Obsidian i Claude Code Pozwól że pokażę jak wygląda mój typowy dzień pracy z tym setup'em. ### Morning routine 1. Otwieram Obsidian + Claude Code 2. Uruchamiam skill do przeglądu dnia - sprawdzam co mam w ClickUp, co w kalendarzu 3. Sprawdzam ważne maile przez skill ze skryptem ### W trakcie pracy - **Quick capture** do inbox w Obsidian - szybkie notatki, pomysły - Proszę Claude o przetworzenie i umieszczenie w odpowiednim folderze - Generuję dokumenty z notatek przez skills - Sync ważnych tasków z ClickUp do notatek projektu ### Research sessions - Definiuję temat który mnie interesuje - Claude Code robi research i zapisuje notatki - Przeglądam i dodaję własne przemyślenia ### Content creation - Notatki → drafty przez skills - Templates zapewniają spójność - Ale **ludzki dotyk** pozostaje kluczowy > "Obsidian to moje płótno. Wszystko co mój second brain generuje - dokumenty, drafty, pomysły - zarządzam tutaj. To tutaj dodaję swój ludzki dotyk." ### Integracje które używam - **ClickUp** - taski i projekty (przez skrypt w skill) - **Gmail/Outlook** - ważne maile do przetworzenia (bezpośrednio przez biblioteki) - **Google Calendar** - spotkania i deadline'y ### Dlaczego skrypty zamiast Make.com/Zapier Make.com pozwala wystawić scenariusze jako narzędzia MCP - to ciekawa opcja. Ale wolę pisać własne skrypty: - Są szybsze - nie ma dodatkowej warstwy komunikacji - Mam pełną kontrolę nad logiką - Wszystko w jednym miejscu, łatwiej debugować ## Jak zacząć - pierwsze kroki Nie musisz budować wszystkiego naraz. Zacznij od podstaw i rozwijaj stopniowo. ### 1. Zainstaluj narzędzia - **Obsidian** - darmowy, pobierz z [obsidian.md](https://obsidian.md) - **Claude Code** - wymaga subskrypcji Claude Pro lub API ### 2. Stwórz strukturę folderów Zacznij od prostej struktury PARA: ```text obsidian-vault/ ├── 00-inbox/ # Wszystko nowe trafia tutaj ├── 01-projects/ # Aktywne projekty ├── 02-areas/ # Stałe obszary odpowiedzialności ├── 03-resources/ # Materiały referencyjne ├── 04-archive/ # Zakończone rzeczy └── .claude/ └── skills/ # Tu będą Twoje skills ``` ### 3. Stwórz pierwszy skill Zacznij prosto - np. skill do podsumowywania notatek: ```yaml --- name: note-summarizer description: Summarize long notes into key points. Use when user asks to summarize or extract key points from notes. --- # Note Summarizer ## Workflow 1. Read the specified note 2. Extract key concepts (5-7 points) 3. Create bullet-point summary 4. Add to top of note or save separately ``` ### 4. Zbuduj workflow - Zdefiniuj rutyny dnia (morning review, end-of-day) - Stwórz templates dla powtarzalnych dokumentów - Testuj i iteruj ### 5. Rozwijaj stopniowo - Dodawaj skills gdy pojawia się realna potrzeba - Nie próbuj budować wszystkiego na start - Każdy skill to inwestycja - upewnij się że będziesz go używać ## Kluczowe wnioski 1. **Claude Code to nie tylko kodowanie** - file operations, search, web search czynią go potężnym asystentem ogólnego przeznaczenia 2. **Obsidian + Claude Code = perfect match** - pliki markdown + lokalne przechowywanie + AI processing to idealna kombinacja 3. **Skills dają context-efficient extensibility** - progressive disclosure zapobiega context bloat 4. **Start simple, grow gradually** - jeden skill na raz, iteruj bazując na realnych potrzebach 5. **Human touch remains essential** - AI augmentuje, nie zastępuje Twojego myślenia 6. **File-based workflow is powerful** - wszystko pod version control, przenośne, prywatne ---

Chcesz zbudować własny system zarządzania wiedzą z AI?

Pomogę Ci zaprojektować i wdrożyć second brain dopasowany do Twoich potrzeb. Od wyboru narzędzi przez konfigurację skills po optymalizację workflow.

Umów bezpłatną konsultację
## Przydatne zasoby - [Obsidian](https://obsidian.md) - oficjalna strona - [Claude Code Documentation](https://docs.anthropic.com/claude-code) - dokumentacja Claude Code - [Cole Medin / Dynamist](https://www.youtube.com/@ColeMedin) - inspiracja dla tego artykułu - [5 technik pracy z Claude Code](/blog/5-technik-pracy-z-claude-code) - powiązany artykuł - [PARA Method](https://fortelabs.com/blog/para/) - system organizacji wiedzy - [second-brain-template](https://github.com/plipowczan/second-brain-template) - gotowy szablon, żeby zacząć budować własny second brain w kilka minut ## FAQ
### Czy potrzebuję umiejętności programowania żeby używać Claude Code jako second brain? Nie, Claude Code obsługuje się przez naturalny język. Wystarczy opisać co chcesz zrobić - "stwórz podsumowanie moich notatek z folderu projekty" - a Claude wykona resztę. Znajomość markdown jest pomocna, ale nie wymagana. Skills piszesz również w markdown, nie w kodzie programistycznym.
### Czym różni się to podejście od używania ChatGPT lub Claude.ai bezpośrednio w przeglądarce? Kluczowa różnica to dostęp do plików lokalnych. Claude Code działa na Twoim komputerze i ma bezpośredni dostęp do plików w Obsidian - może je czytać, edytować, tworzyć nowe. W przeglądarce musisz ręcznie kopiować treść. Dodatkowo skills pozwalają na automatyzację powtarzalnych workflow bez utraty kontekstu między sesjami.
### Jak połączyć second brain z innymi narzędziami jak email czy task manager? Masz dwa podejścia: MCP servers (bezpośrednie połączenie z Gmail, Outlook, ClickUp przez biblioteki lub serwery MCP) lub skrypty w skills (moje preferowane). Skrypty w Python/Node.js dodajesz do folderu skills i Claude Code je uruchamia. Make.com pozwala wystawić scenariusze jako narzędzia MCP, ale wolę skrypty - są szybsze, prostsze do debugowania i wszystko mam w jednym miejscu.
### Czy moje notatki są bezpieczne i prywatne przy używaniu Claude Code? Tak, notatki pozostają na Twoim komputerze - Obsidian nie wymaga chmury. Claude Code przetwarza pliki lokalnie i wysyła do API tylko to, co jest potrzebne do danego zadania. Nie musisz synchronizować całego vault'a z zewnętrznym serwisem jak w przypadku Notion. Masz pełną kontrolę nad swoimi danymi.
### Od czego najlepiej zacząć jeśli nigdy nie używałem Obsidian ani Claude Code? Zacznij od zainstalowania Obsidian i stworzenia prostej struktury folderów (inbox, projekty, zasoby). Następnie zainstaluj Claude Code i przetestuj podstawowe operacje - poproś o podsumowanie pliku, stworzenie nowej notatki. Dopiero gdy poczujesz się komfortowo, dodaj pierwszy prosty skill. Cały proces można rozłożyć na tydzień nauki po 30 minut dziennie.
### Czy mogę używać tego systemu z innymi edytorami markdown zamiast Obsidian? Tak, Claude Code działa z dowolnymi plikami markdown. Obsidian jest rekomendowany ze względu na graph view (wizualizacja połączeń), wbudowane linkowanie i rozbudowany ekosystem pluginów. Alternatywy jak Logseq czy Foam również zadziałają, ale mogą wymagać dostosowania workflow. Kluczowe jest używanie lokalnych plików markdown, nie aplikacji chmurowych.
### Ile kosztuje taki setup i jakie są wymagania sprzętowe? Obsidian jest darmowy do użytku osobistego. Claude Code wymaga subskrypcji Claude Pro ($20/miesiąc) lub dostępu przez API. Wymagania sprzętowe są minimalne - każdy współczesny komputer z 8GB RAM wystarczy. Vault Obsidian może mieć tysiące notatek bez problemów z wydajnością.
--- # 15 hacków do Cursor.sh które zmienią sposób pracy z AI Source: https://pawel.lipowczan.pl/blog/15-cursor-hacks-produktywnosc-ai Published: 2026-01-15 Przez pierwsze trzy miesiące używałem Cursor jak zwykły VS Code z autocompletem. Płaciłem **$20 miesięcznie**, żeby agent odpisywał "sure, let me help you with that" i generował kod który i tak musiałem przepisać. Brzmi znajomo? Dopiero gdy zacząłem zagłębiać się w ukryte funkcje, zrozumiałem że większość użytkowników - ja włącznie - wykorzystuje **zaledwie 20% możliwości** tego narzędzia. To jak kupić iPhone'a i używać go tylko do dzwonienia. Problem nie leży w Cursor. Leży w tym, że **najlepsze funkcje są ukryte**, nieintuicyjne, albo schowane głęboko w ustawieniach. A każdy dzień bez ich znajomości to stracone pieniądze, zmarnowany **context window** i frustracja. W tym artykule pokażę Ci **16 hacków** (tak, będzie bonus) które zmieniły sposób w jaki pracuję z AI. Od podstawowych skrótów klawiszowych, które oszczędzą Ci godziny klikania, przez zaawansowane techniki jak **worktrees** do równoległego testowania wielu modeli, aż po **strukturyzację promptów** która sprawia że 80% kodu z pierwszej iteracji idzie prosto do produkcji. Podzieliłem je na trzy poziomy trudności, więc możesz zacząć od podstaw i stopniowo odkrywać kolejne warstwy. Gotowy żeby wycisnąć 100% z subskrypcji Cursor? ## Dlaczego warto znać Cursor w 100% Zanim przejdziemy do hacków, porozmawiajmy o słoniu w pokoju: **czy naprawdę warto się w to zagłębiać?** **Context window** to najcenniejsza nieruchomość w świecie AI. Za każdym razem gdy włączysz niepotrzebne MCP, za każdym razem gdy agent "zapomina" co robił 20 wiadomości temu, za każdym razem gdy przekraczasz limit subskrypcji w środku sprintu - to nie bug, to twoja niewiedza jak zarządzać tym zasobem. Za **$20 miesięcznie** (niecałe 80 zł) możesz mieć drogi autocomplete. Albo **3-5x boost produktywności** który zwraca się w pierwszy dzień miesiąca. Różnica? Wiedza o ukrytych funkcjach. Pierwsze trzy miesiące używałem Cursor jak amator. Klikałem myszką żeby otworzyć terminal. Czekałem bezczynnie aż agent skończy, sprawdzając Instagram. Przekraczałem limity, bo nie monitorowałem zużycia. Tracąc czas na powtarzalne prompty, bo nie wiedziałem o custom commands. Odkrywając kolejne funkcje zrozumiałem że Cursor to nie tylko editor - to **ekosystem narzędzi**. Które - gdy użyte prawidłowo - integrują się z Claude Code, worktrees, dokumentacją bibliotek i własnymi workflow. To daje przewagę której reszta po prostu nie ma. ## Poziom 1 - Podstawy (Każdy powinien znać) Te hacki powinien znać każdy użytkownik Cursor, niezależnie od poziomu zaawansowania. Jeśli nie znasz tych rzeczy, tracisz czas i pieniądze każdego dnia. ### Hack 1: Skróty klawiszowe które zaoszczędzą godziny Przez miesiąc klikałem myszką w terminal. Serio. Za każdym razem gdy chciałem sprawdzić output. Potem odkryłem `Cmd+J` i poczułem się jak idiota. Skróty klawiszowe to najszybszy sposób żeby przestać walczyć z interfejsem i zacząć walczyć z problemami. Te podstawowe skróty oszczędzą Ci dosłownie **godziny klików tygodniowo**: ```text Podstawowe skróty: Cmd+B (Ctrl+B) - Sidebar on/off Cmd+J (Ctrl+J) - Terminal on/off Cmd+E (Ctrl+E) - Przełącz na agenta Shift+Tab - Zmień tryb (Ask/Agent/Plan/Debug) Cmd+/ (Ctrl+/) - Wybierz model ``` Najważniejsze: **`Cmd+E`** przełącza Cię między edytorem a agentem natychmiast. Nie szukasz kursorem, nie klikasz. Piszesz kod → `Cmd+E` → dajesz agentowi zadanie → `Cmd+E` → wracasz do kodu. **`Shift+Tab`** pozwala Ci cyklicznie przełączać między trybami Ask, Agent, Plan i Debug. Większość użytkowników klika dropdown. Ty będziesz robił to w ułamku sekundy. Zainwestuj 10 minut żeby nauczyć się tych skrótów. Zwróci Ci się to w pierwszy dzień. ⚠️ Protip: Możesz sobie pomóc za pomocą urządzenia takiego jak [Stream Deck](https://www.elgato.com/us/en/p/stream-deck). Możesz na nim stworzyć przyciski dla każdego z trybów i używać ich jako skrótów klawiszowych. Mnie to bardzo pomaga nie tylko przy wykorzystaniu skrótów, ale także przy ich nauce. ### Hack 2: Pokaż zużycie subskrypcji - nie daj się zaskoczyć limitem Raz przekroczyłem limit w środku tygodniowego sprintu. Cursor przestał działać. Deadline nie poczekał. Nigdy więcej. Domyślnie Cursor pokazuje **zużycie subskrypcji dopiero gdy jesteś blisko limitu**. To za późno. Nie możesz zarządzać tym czego nie widzisz. Rozwiązanie: włącz stałe wyświetlanie usage summary. ```text Ścieżka w ustawieniach: Settings → Agents → Usage Summary → "Always" ``` Teraz na dole interfejsu zawsze widzisz ile kredytów zużyłeś. Możesz **świadomie decydować** kiedy używać Opus (drogi, ale genialny), a kiedy przełączyć się na Sonnet (szybki, tańszy). Monitoruj to regularnie. Gdy zbliżasz się do **60-70% limitu** w połowie miesiąca, to sygnał że musisz: - Przełączyć się na lżejsze modele - Wyłączyć niepotrzebne MCP (o tym za chwilę) - Restart konwersacji zamiast dokładania wiadomości To prosta zmiana która może uratować Ci cały sprint. ### Hack 3: Włącz dźwięki zakończenia - przestań czekać bezczynnie Ile razy sprawdziłeś telefon podczas gdy agent skończył 5 minut temu? Ile razy przełączyłeś się na Instagrama "na chwilę" i wróciłeś po 15 minutach? Za dużo. Problem nie jest w Tobie. Problem jest w tym że **Cursor nie mówi Ci gdy skończy**. Siedzisz, czekasz, patrzysz. Albo robisz coś innego i gubisz momentum. Rozwiązanie: włącz **completion sound**. ```text Settings → General → Completion sound (włącz) ``` Teraz gdy agent skończy zadanie, usłyszysz subtelny dźwięk. Możesz sprawdzać dokumentację, pisać notatki, nawet zrobić kawę - i wrócisz dokładnie gdy Cursor będzie gotowy. To brzmi jak detal, ale w praktyce **zmienia sposób pracy**. Zamiast bezczynnie czekać, wykorzystujesz czas. Zamiast przełączać się na rozpraszacze (Instagram, Twitter), zostanjesz w flow. Drobna zmiana, ogromna różnica w produktywności. ### Hack 4: Early Access - dostęp do nowych funkcji miesiące wcześniej Custom modes pojawiły się w **Early Access** 2 miesiące przed oficjalną wersją. W tym czasie użytkownicy którzy wiedzieli o tej opcji mieli dostęp do funkcji o których inni nawet nie słyszeli. Większość użytkowników siedzi na domyślnej wersji Cursor. Czekają miesiącami na nowe funkcje, które już są dostępne - tylko ukryte za jednym przełącznikiem. ```text Settings → Beta → Update Access → Early Access ``` Dostępne opcje: - **Default** - stabilna wersja, opóźnione updates - **Early Access** - nowe funkcje miesiące wcześniej, zwykle stabilne - **Nightly** - dla deweloperów, może być buggy Z mojego doświadczenia, **Early Access to najlepsza opcja**. Dostajesz nowe funkcje znacznie wcześniej, a stabilność jest w 95% przypadków idealna. Jedyne co ryzykujesz to okazjonalny bug - który zwykle jest naprawiony w ciągu dni. Dlaczego większość nie wie o tej opcji? Bo jest schowana w Beta settings. Bo Cursor nie krzyczy o tym. Bo zakłada się że jesteś zadowolony z domyślnych ustawień. Nie bądź zadowolony z domyślnych ustawień. Przełącz się na Early Access i bądź o krok przed resztą. ### Hack 5: Wiele okien - pracuj nad dwoma projektami jednocześnie To nie jest odkrywczy hack - większość editorów ma tę funkcję. Ale wiele osób nie zdaje sobie sprawy że mogą mieć **kilka projektów otwartych jednocześnie** z niezależnymi konwersacjami AI. **File → New Window** otwiera nową instancję Cursor. Każda z własnym projektem, własnym agentem, własnym context window. **Use cases:** - Referencujesz stary projekt podczas budowy nowego - Kopiujesz pattern z jednego projektu do drugiego bez przełączania kontekstu - Pracujesz nad wieloma klientami jednocześnie - Porównujesz implementacje między projektami **Uwaga:** Każde okno = osobna instancja w pamięci. Jeśli masz **8GB RAM**, dwa okna mogą być już na granicy. Z **16GB+** możesz swobodnie otworzyć 3-4 projekty. Nie robi to takiej różnicy jak worktrees (o tym później). Ale przydatne gdy pracujesz nad wieloma projektami. Szczególnie gdy chcesz **skopiować pattern** z jednego projektu do drugiego bez przełączania kontekstu i gubienia wątku konwersacji z agentem. Proste, skuteczne, niedoceniane. ## Poziom 2 - Średniozaawansowany Tutaj zaczyna się magia. Te techniki odróżniają tych, którzy znają Cursor, od tych którzy go **opanowali**. ### Hack 6: Zarządzaj MCP - nie marnuj context window Raz miałem **6 MCP włączonych** jednocześnie. Context window wypalał się w 10 wiadomości. Agent "zapominał" co robiliśmy na początku konwersacji. Musiałem restartować co 15 minut. **MCP** (Model Context Protocols) to integracje które rozszerzają możliwości agenta. Browser automation, terminal access, file operations - każde MCP dodaje nowe supermoce. Problem? **Każde aktywne MCP zjada ogromny kawałek context window**. Nawet gdy go nie używasz. Po prostu swoją obecnością. Strategia: ```text Rekomendowany setup MCP: ✅ Browser Automation (często przydatne) ❌ Pozostałe (włącz tylko gdy potrzeba) Ścieżka: Settings → Tools & MCP → [MCP Name] (wyłącz) ``` **Default approach (błędny):** Włącz wszystkie MCP "na wszelki wypadek". Miej pełną paletę narzędzi. **Pro approach:** **Disable all** na starcie projektu. Włącz konkretne MCP dopiero gdy faktycznie je potrzebujesz. Agent powie Ci gdy będzie czegoś potrzebował. Praktycznie jedyne MCP które zostawiam włączone zawsze to **Browser Automation** - bo często pracuję z UI testing i automatyzacją. Reszta? Włączam on-demand. Ta zmiana w podejściu **zwiększyła długość moich konwersacji z agentem 2-3x**. Mniej restartów, lepsza pamięć kontekstu, płynniejszy flow. ### Hack 7: Własne komendy - przestań powtarzać te same prompty Mam **8 custom commands**. Najczęściej używam `/prime` (ładuje kontekst projektu) i `/package-health` (sprawdza dependencies). Wcześniej wpisywałem te same prompty ręcznie. Teraz? Dwa znaki. **Slash commands** to reusable prompt templates. Piszesz raz, używasz setki razy. **Jak stworzyć:** 1. Wpisz `/` w agencie 2. Wybierz **"Create command"** 3. Napisz prompt template 4. Zapisz w `.cursor/commands/` lub `.claude/commands/` Przykład: ```bash # Przykład custom command: Package Health Check # Plik: .cursor/commands/package-health.md Scan node_modules and package.json for: - Security vulnerabilities - Outdated dependencies - Breaking changes in recent versions Report findings with severity levels. ``` Teraz wpisujesz `/package-health` i agent dokładnie wie co zrobić. Żadnych wyjaśnień, żadnego kontekstu - sam wykonuje rutynową inspekcję. **Pro tip:** Możesz **importować commands z Claude Code**. ```text Settings → Rules and Commands → Import Claude commands ``` Claude Code ma gotowe komendy które możesz zaadaptować. Nie wymyślaj koła na nowo. **Kiedy stworzyć komendę?** Jeśli robisz coś więcej niż 2x, zamiast kopiować prompt - stwórz komendę. Proste. Efektywne. Game-changing. ### Hack 8: Zarządzanie kontekstem - reguła 60% Powyżej **60% zapełnienia context window** agent zaczyna "zapominać" wcześniejszych instrukcji. To nie subiektywne wrażenie - to empirycznie sprawdzone. AI świetnie pamięta **początek i koniec** konwersacji. Słabo pamięta **środek**. To tzw. "lost in the middle" problem. Cursor pokazuje **context window indicator** poniżej pola tekstowego wiadomości, obok opcji agenta. Monitoruj go jak jastrząb: ```text Strategia context management: < 40% - Kontynuuj swobodnie 40-60% - Szukaj naturalnego breakpoint > 60% - Rozważ restart (nowa konwersacja) > 80% - Definitywnie restart ``` **Naturalny breakpoint** to moment gdy: - Skończyłeś feature i commitowałeś - Przełączasz się na inny obszar projektu - Agent wykonał zadanie i jest "gotowy" na nowe wyzwanie **Dwie strategie restartu:** 1. **Summary approach:** Poproś agenta o podsumowanie konwersacji, skopiuj do nowej. Dobre gdy masz skomplikowany kontekst który chcesz przenieść. 2. **Fresh start:** Po prostu zacznij nową konwersację. Dobre gdy kontekst był specyficzny dla jednego zadania i nie będzie potrzebny. Ja używam głównie **fresh start**. Mniej overhead, czysty umysł, agent bez bagażu z przeszłości. Pamiętaj: **60% to soft limit**. Możesz iść wyżej, ale jakość odpowiedzi zaczyna spadać. 80% to hard limit - poza tym wszystko się sypie. ### Hack 9: Indeksuj dokumentację - pełna wiedza o bibliotekach w Cursor **Tailwind CSS v4** wyszedł kilka miesięcy temu. Agent ciągle mylił składnię, proponował przestarzałe klasy. Frustracja level max. Zaindeksowałem oficjalną dokumentację Tailwind v4 w Cursor. Problem zniknął. Agent znał nową składnię lepiej niż ja. **Dodawanie dokumentacji:** ```text 1. Znajdź official docs URL (np. https://tailwindcss.com/docs) 2. Settings → Indexing and Docs → Add Doc 3. Wklej URL → Confirm 4. Użyj: @ → Docs → [Nazwa biblioteki] ``` Cursor **zeskrapuje i indeksuje całą dokumentację**. Przykład: Tailwind docs to **439 stron** które agent ma w pamięci. Za każdym razem gdy użyjesz `@ Docs → Tailwind CSS`, agent ma dostęp do pełnej, aktualnej dokumentacji. **Dlaczego to robi różnicę:** - **Cutoff window problem solved:** Agent nie jest ograniczony do wiedzy sprzed stycznia 2025 - **Aktualne API:** Nowe frameworki, nowe wersje bibliotek - agent zna aktualną składnię - **Zero halucynacji:** Zamiast "wymyślać" API, czyta z oficjalnej dokumentacji **Pro tip:** Zaindeksuj dokumentację bibliotek których używasz regularnie. Dla mnie to: - Tailwind CSS - React (szczególnie nowe hooki) - Convex - Framer Motion Agent przestaje być "ogólnym AI" i staje się **ekspertem w Twoim stack'u**. ### Hack 10: Agent steering - przejmij kontrolę w locie Wyobraź sobie: agent pisze komponent. W połowie realizujesz że chcesz dodać jeszcze jedną funkcję. Czekasz aż skończy? Przerywasz i zaczynasz od nowa? Są lepsze opcje. **Agent steering** to możliwość wysyłania wiadomości gdy agent pracuje. Albo czekasz aż skończy, albo przerywasz natychmiast. **Gdzie znaleźć:** ```text Lokalizacja ustawienia: Agent pane → ... (trzy kropki) → Agent Settings → Queue Messages Dostępne opcje: 1. Send after current message → Agent kończy obecne zadanie, potem wykonuje nowe → Idealne dla dodawania zadań do kolejki 2. Stop & send right away → Natychmiastowe przerwanie i nowe zadanie → Używaj gdy kierunek pracy wymaga zmiany ``` **"Send after current message"** używam **najczęściej**. Agent kończy obecne zadanie (np. komponent), potem automatycznie przechodzi do kolejnego z listy (np. testy do tego komponentu). Efektywny kolejkowanie bez mikro-managementu. **"Stop & send right away"** używam **tylko** gdy: - Agent poszedł w złym kierunku i każda sekunda to strata - Zmieniły się wymagania i obecne zadanie już nie ma sensu - Potrzebuję natychmiastowej odpowiedzi na pytanie **Caveat:** Nadużywanie "Stop & send" powoduje chaos. Agent gubił się w kontekście. Frustracja rośnie. Używaj rzadko i tylko gdy naprawdę musisz zmienić kierunek. **Pro workflow:** Planujesz listę zadań. Pierwszą wysyłasz normalnie. Kolejne queue'ujesz przez "Send after current message". Agent pracuje jak maszyna przez kolejne zadania, a Ty możesz przełączyć się na coś innego. ## Poziom 3 - Pro Features To są funkcje, o których większość użytkowników nie wie że istnieją. Ale profesjonaliści używają ich **codziennie**. ### Hack 11: Worktrees - testuj wiele rozwiązań równolegle **Worktrees** to funkcja która brzmi skomplikowanie, ale zmienia wszystko. Pozwala Ci **testować 3 różne podejścia jednocześnie** bez czekania na kolejne iteracje. **Czym są worktrees:** **Git worktrees** to izolowane kopie projektu na osobnych branchach. Cursor wykorzystuje ten mechanizm do równoległego testowania wielu rozwiązań. Każda wersja pracuje w swojej kopii kodu, bez konfliktów. Po zakończeniu wybierasz najlepsze rozwiązanie i przenosisz do głównego brancha. **Jak to działa krok po kroku:** 1. **Start:** Otwierasz chat z agentem w Cursor 2. **Wybór trybu:** Poniżej pola tekstowego widzisz opcje: **Local / Worktree / Cloud** - wybierasz **Worktree** 3. **Wybór modelu:** Klikasz w pole wyboru modelu 4. **Multiple models lub mnożnik:** Teraz masz dwie opcje: - **Multiple models:** Zaznacz "use multiple models" i wybierz różne modele (np. Composer + Sonnet + GPT-4o) - **Mnożnik:** Wybierz jeden model i ustaw mnożnik 2x, 3x lub 4x (np. 3x Claude Sonnet - dostaniesz 3 różne wersje od tego samego modelu) 5. **Automatyczny setup:** Cursor tworzy git worktree dla każdej wersji, kopiuje projekt, uruchamia serwery na osobnych portach 6. **Równoległa praca:** Wszystkie wersje pracują jednocześnie nad tym samym zadaniem 7. **Preview:** Sprawdzasz rezultaty na różnych portach (localhost:3001, 3002, 3003) 8. **Apply:** Wybierasz najlepsze rozwiązanie → Apply → zmiany trafiają do głównego brancha **Setup trick - autostart serwerów:** ```json // .cursor/worktrees.json - automatyczny setup // Umieść w głównym katalogu projektu { "commands": [ "npm install", // Instaluje dependencies w worktree "npm run dev" // Uruchamia dev server automatycznie ] } ``` **Praktyczny workflow:** ```bash # Praktyczny przykład użycia worktrees: # 1. Masz zadanie: "Add dark mode toggle with smooth transition" # 2. Chat z agentem → wybierz "Worktree" (zamiast Local/Cloud) # 3. Kliknij wybór modelu → zaznacz "use multiple models" # 4. Wybierz: Composer + Sonnet + GPT-4o (lub ustaw 3x Sonnet dla różnorodności) # 5. Cursor automatycznie: # - Tworzy branch-wt-composer-xyz # - Tworzy branch-wt-sonnet-abc # - Tworzy branch-wt-gpt4o-def # - Uruchamia npm install + npm run dev w każdym # 6. Po 2-3 minuty masz 3 działające wersje: # - localhost:3001 (Composer) - minimalistyczny przełącznik # - localhost:3002 (Sonnet) - animowany slider z ikonkami # - localhost:3003 (GPT-4o) - toggle z preview kolorów # 7. Otwierasz wszystkie 3 w przeglądarce, porównujesz # 8. Sonnet wygląda najlepiej → Review changes → Apply # 9. Zmiany trafiają do głównego brancha, worktrees są czyszczone ``` **Dlaczego to robi różnicę:** - **Design choices:** 3 modele = 3 różne podejścia wizualne, wybierasz najlepsze - **Refactoring:** różne strategie implementacji, porównujesz jakość kodu - **Bug fixes:** widzisz które podejście jest najbardziej eleganckie - **Zero strat czasu:** Bez worktrees czekałbyś 15 minut na 3 iteracje. Z worktrees? 5 minut i masz wszystkie wersje jednocześnie. **Real use case z mojego workflow:** Design changes - puszczam 3 modele jednocześnie. W 5 minut widzę 3 różne podejścia wizualne. Bez worktrees musiałbym czekać 15 minut na kolejne iteracje, a każda następna byłaby "kontaminowana" feedbackiem z poprzedniej. **Kiedy NIE używać worktrees:** - Proste bugfixy (overkill) - Jasne wymagania (jeden model wystarczy) - Ograniczone kredyty (3 modele = 3x koszt) Worktrees to bazooka. Używaj tylko gdy potrzebujesz bazooki 😉 ### Hack 12: Dwu-modelowy workflow - GPT-5.2 planuje, Claude wykonuje **GPT-5.2 High** to najlepszy model do planowania. Tworzy szczegółowe, przemyślane plany które uwzględniają edge case'y i architekturę. Jest tylko jeden problem: jest **wolny jak cholera** przy implementacji. **Claude Sonnet** jest szybki przy kodowaniu. Decent przy planowaniu, ale nie na poziomie GPT-5. Rozwiązanie? **Użyj obu**. ```text Optymalny workflow: 1. Plan Mode → GPT-5.2 High (plan) 2. Po planie: Przełącz model → Claude 4.5 Sonnet 3. Kliknij "Build" Rezultat: Najlepszy plan + szybka implementacja ``` **Jak to działa:** 1. Dajesz zadanie w **Plan Mode** 2. Wybierasz **GPT-5.2 High** jako model 3. Czekasz (tak, to trwa) aż GPT stworzy szczegółowy plan 4. **ZANIM klikniesz "Build"** - przełączasz model na **Claude 4.5 Sonnet** 5. Teraz klikasz "Build" 6. Claude implementuje plan GPT - szybko i efektywnie **Rezultat:** Dostajesz **dokładność planowania GPT** + **szybkość implementacji Claude**. Oszczędzam tym trickiem **70% czasu** w porównaniu do używania samego GPT do build, albo samego Claude do planu. **Caveat:** GPT-5.2 zżera sporo kredytów. Używaj tej techniki do **większych features**, nie do mikro-tasków. Dla prostych rzeczy Claude sam wystarczy. To najlepszy przykład tego że **wybór modelu nie musi być decyzją "albo-albo"**. Możesz łączyć silne strony różnych modeli w jednym workflow. ### Hack 13: Strukturyzacja promptów - user stories i design patterns **Problem:** Większość developerów pisze prompty ad-hoc. "Add login". "Fix this bug". "Make it prettier". Agent dostaje niejasne wymagania i generuje generyczny kod. Zanim odkryłem strukturyzację promptów, **50% pierwszych iteracji z AI było do wyrzucenia**. Kod działał technicznie, ale nie spełniał wymagań. Bo wymagania były niejasne. **Rozwiązanie:** Strukturyzuj prompty według wzorców, których AI są trenowane - **user stories** i **design patterns**. **User Story Format (dla feature development):** ```text As a [user type] I want [goal/desire] So that [benefit/value] Acceptance Criteria: - [ ] Criterion 1 - [ ] Criterion 2 - [ ] Criterion 3 Technical Context: - Existing components: [list] - Design system: [colors, spacing] - Similar implementations: [reference files] ``` **Dlaczego to działa:** - AI są **trenowane na user stories** z GitHub/Jira - to ich native language - Format **wymusza jasność wymagań** - nie możesz być vague - **Acceptance criteria = built-in testing checklist** - agent wie kiedy skończył - **Technical context eliminuje halucynacje** - agent referencuje istniejący kod **Design Pattern Prompts (dla architektury):** ```text I need to implement [feature] using [pattern name] pattern. Pattern Structure: - Component A: [responsibility] - Component B: [responsibility] - Communication: [how they interact] Constraints: - Must work with existing [X] - Performance requirement: [Y] - Should follow project convention: [Z] References in codebase: - Similar pattern used in: [file path] ``` **Real Examples - Bad vs Good:** **Bad prompt:** ```text Add user authentication ``` **Good prompt (user story):** ```text As a returning user I want to log in with email/password So that I can access my saved preferences Acceptance Criteria: - [ ] Login form with email + password fields - [ ] Form validation (email format, password min 8 chars) - [ ] Error handling for invalid credentials - [ ] Success: redirect to dashboard - [ ] Failure: show error message, keep user on page Technical Context: - Auth provider: Clerk (already configured) - Similar form: src/components/ContactForm.jsx - Design system: Tailwind with dark-800 backgrounds - State management: React Context in src/context/AuthContext.jsx ``` Widzisz różnicę? **Bad prompt** to zgadywanka. Agent musi założyć 10 rzeczy. **Good prompt** to instrukcja obsługi. Agent wie dokładnie co zrobić, jak, i co znaczy "skończone". Teraz **80% kodu z pierwszej iteracji idzie do produkcji**. Nie dlatego że AI stało się lepsze. Dlatego że ja stałem się lepszy w komunikacji z AI. **Bonusowy trick - Image-to-code prompts:** ```text [Załącz screenshot/mockup] Recreate this design with following specifications: - Framework: React + Tailwind - Components to extract: [list] - Interactive elements: [behaviors] - Responsive breakpoints: mobile (< 768px), desktop (>= 768px) - Color palette: [specify if different from image] DO NOT: [lista czego unikać] ``` Agent dostaje visual reference + technical context. Rezultat? Pixel-perfect implementation z pierwszej próby. **Kiedy używać strukturyzacji:** - ✅ Nowe features (user stories) - ✅ Refactoring (design patterns) - ✅ Design implementation (image-to-code) - ❌ Proste bugfixy (overkill) - ❌ Eksploracyjne zadania (za dużo strukty krępuje kreatywność) ### Hack 14: @ command power features - kontekst na wyciągnięcie ręki Zamiast opisywać słowami "ten plik z auth", wpisuję `@ → Files → AuthContext.jsx` i agent od razu ma pełny kontekst. Zero friction. **@ command** to szybki sposób dodawania kontekstu do agenta. Nie musisz copy-paste kodu, nie musisz opisywać - po prostu wskazujesz. **Dostępne opcje:** ```text @ command - dostępne opcje: 1. Files & Folders → Dodaj konkretne pliki do kontekstu → Przykład: @ Files → src/components/Header.jsx 2. Docs → Zindeksowana dokumentacja bibliotek → Przykład: @ Docs → Tailwind CSS 3. Terminals → Output z terminala (błędy, logi) → Przydatne przy debugowaniu 4. Branch (Diff with main) → Różnice między branchami → Zobacz co się zmieniło w feature branchu → Przykład: "Co dodałem w tym branchu?" 5. Browser → Zawartość otwartej strony → Design inspiration, dokumentacja online Wszystkie opcje można łączyć dla pełnego kontekstu! ``` **Use cases:** - **Files:** "Sprawdź ten plik zanim zasugerujesz zmiany" → `@ Files → auth.js` - **Docs:** "Użyj oficjalnej dokumentacji Tailwind v4" → `@ Docs → Tailwind CSS` - **Branch:** "Pokaż co się zmieniło od main brancha" → `@ Branch (Diff with main)` - **Browser:** "Zaimplementuj design z tej strony" → `@ Browser` **Pro tip:** Możesz dodać **wiele kontekstów naraz**: "Zaimplementuj authentication flow similar to auth.js, using Clerk docs, based on the changes in current branch" `@ Files → auth.js` + `@ Docs → Clerk` + `@ Branch` Agent ma **full picture**. Nie zgaduje, nie halucynuje. Po prostu implementuje zgodnie z tym co widzi. ### Hack 15: Duplicate - klonuj dobrze przygotowany kontekst Przez **rok** używałem Cursor i nie wiedziałem że to istnieje. Schowane głęboko w widoku zarządzania agentami. **Use case:** Spędziłeś 30 minut na "primingu" agenta. Załadowałeś docs, rules, kontekst projektu. Agent rozumie codebase idealnie. Teraz chcesz zaimplementować **3 powiązane features** z tym samym fundamentem. Opcja A: Re-prime agenta 3 razy. 90 minut zmarnowanych. Opcja B: **Duplicate**. 3 kliki. **Gdzie znaleźć:** ```text Workflow z duplicate: 1. Prime agent: załaduj docs, rules, kontekst projektu 2. Gdy agent dobrze rozumie codebase 3. Przejdź do widoku zarządzania agentami (lista wszystkich agentów) 4. Znajdź agenta którego chcesz sklonować 5. Kliknij ... (trzy kropki) obok tego agenta 6. Wybierz "Duplicate" 7. Nowy agent z identycznym kontekstem 8. Implementuj kolejną feature bez powtarzania primingu ⚠️ Jeśli duplikat nie działa: restart Cursor ``` **UWAGA:** NIE znajdziesz tej opcji w menu `...` bezpośrednio w agent pane podczas rozmowy. Musisz przejść do **widoku zarządzania agentami** (lista wszystkich agentów w sidebarze). ### BONUS Hack 16: .cursorignore manipulation - niebezpieczny ale przydatny ⚠️ **OSTRZEŻENIE: Używaj TYLKO w środowisku testowym. NIGDY w produkcji.** `.cursorignore` kontroluje co AI może widzieć. Domyślnie blokuje pliki `.env` i `.env.local` dla bezpieczeństwa. I słusznie - **nie chcesz** żeby API keys trafiły do konwersacji z AI. Ale czasami - w **testowym repo z dummy credentials** - chcesz żeby agent mógł **zwalidować setup zmiennych środowiskowych**. **Hack:** ```bash # .cursorignore file # ⚠️ UWAGA: Używaj TYLKO w środowisku testowym # Default (bezpieczne): .env .env.local # Hack - odblokowanie (TYLKO DLA TESTÓW): !.env.local # Pozwala AI widzieć zmienne środowiskowe # Użyj tylko z dummy/test credentials! ``` Prefiks `!` neguje regułę. Agent teraz widzi `.env.local`. **Use case:** Agent może sprawdzić czy masz wszystkie potrzebne zmienne, czy nazwy są poprawne, czy wartości mają odpowiedni format (oczywiście nie widzi samych wartości produkcyjnych - bo to TEST REPO). **Personal warning:** Nigdy nie rób tego w produkcji. Tylko w testowym repo gdzie klucze są dummy. Jeśli masz wątpliwości - **nie rób tego w ogóle**. Pokazuję ten hack dla kompletności. Ale używaj z rozwagą. ## Bonus Tips - Szybkie Wskazówki Te nie zmieściły się w Top 15, ale używam ich regularnie: - **Split terminals:** Hover nad terminalem, kliknij ikonę split. Dwa terminale obok siebie - np. jeden dla dev server, drugi dla git commands. - **Auto theme switching:** Editor Settings → Auto detect color scheme. Cursor przełącza się między jasnym/ciemnym motywem razem z systemem. Drobny detail, ale wygodny. - **Generate cursor rules from docs:** Slash command który wyciąga rules bezpośrednio z dokumentacji projektu. Zamiast pisać ręcznie, agent sam buduje `.cursorrules` na bazie README i CONTRIBUTING. - **Design mode for prototyping:** Custom command który każe agentowi używać tylko mock data. Idealne do prototypowania UI bez backendowej infrastruktury. - **Agent review:** Automated code review on commit. Agent sprawdza kod przed commitem i flaguje potencjalne problemy. Działa jak pre-commit hook z AI. Każdy z tych tricks to oszczędność kilku minut dziennie. Razem - godziny tygodniowo. ## Kluczowe wnioski Po przejściu przez wszystkie 16 hacków, czas na podsumowanie tego co naprawdę ma znaczenie: 1. **Skróty klawiszowe oszczędzają godziny** - `Cmd+J`, `Cmd+E`, `Shift+Tab` to absolutna podstawa. Przestań klikać, zacznij używać klawiatury. 2. **Context window to cenny zasób** - Zarządzaj MCP, resetuj konwersację po 60%, bądź świadomy kosztów. Każda niewykorzystana wiadomość to zmarnowane pieniądze. 3. **Custom commands eliminują powtarzanie** - Jeśli robisz coś więcej niż 2x, stwórz komendę. Proste jak konstrukcja cepa. 4. **Dokumentacja w Cursor = superpowers** - Zaindeksuj biblioteki których używasz. Agent przestaje zgadywać i zaczyna **wiedzieć**. 5. **Strukturyzuj prompty jak profesjonaliści** - User stories i design patterns = 80% kodu z pierwszej iteracji do produkcji. Największa zmiana w moim workflow. 6. **Advanced features nie są intuicyjne** - Worktrees, strukturyzacja promptów, duplicate chat są **ukryte**, ale game-changing. Musisz ich świadomie szukać. 7. **ROI z subskrypcji rośnie z wiedzą** - **$20/miesiąc** to mało lub dużo w zależności od tego jak używasz. Po zastosowaniu tych hacków, zwrot pojawia się pierwszego dnia miesiąca. ## Następne kroki Nie próbuj wdrożyć wszystkiego naraz. To droga do porażki. Zamiast tego: **Dzisiaj (10 minut):** - Włącz **usage summary** (Settings → Chat → Always) - Włącz **early access** (Settings → Beta) - Włącz **completion sounds** (Settings → General) **Ten tydzień (1 godzina):** - Naucz się skrótów: `Cmd+J`, `Cmd+E`, `Shift+Tab` - Sprawdź które **MCP** masz włączone, wyłącz niepotrzebne - Zacznij monitorować **context window** - resetuj po 60% **Ten miesiąc (3-4 godziny):** - Stwórz pierwszą **custom komendę** dla powtarzalnego zadania - Zaindeksuj dokumentację głównej biblioteki z Twojego stacku - Zacznij strukturyzować prompty w formacie **user stories** **Za miesiąc (eksperymentuj):** - Wypróbuj **worktrees** przy następnej design decision - Przetestuj **dwu-modelowy workflow** (GPT plan + Claude build) - Zagłęb się w **advanced prompt patterns** Te hacki zmieniły sposób w jaki pracuję z AI. Zaczynałem od podstaw i stopniowo odkrywałem kolejne warstwy. **Ty możesz przejść tę drogę szybciej** - bo masz ten przewodnik. Powodzenia. I pamiętaj - **$20 miesięcznie** to inwestycja, nie koszt. Pod warunkiem że wiesz jak to wykorzystać. ---

Chcesz maksymalnie wykorzystać AI w kodowaniu?

Pomogę Ci zoptymalizować workflow z AI coding assistants, zbudować custom automation i szkolenia dla zespołu. Od strategii przez implementację po advanced techniques.

Umów bezpłatną konsultację
## Przydatne zasoby - [Cursor Documentation](https://cursor.com/docs) - Oficjalna dokumentacja - [5 technik pracy z Claude Code](/blog/5-technik-pracy-z-claude-code) - Komplementarny artykuł o Claude Code - [Cursor Community Discord](https://discord.gg/cursor) - Społeczność użytkowników - [GitHub: claude-piv-skeleton](https://github.com/plipowczan/claude-piv-skeleton) - Workflow methodology z custom commands ## FAQ
### Czy subskrypcja Cursor za $20 miesięcznie opłaca się początkującym programistom i jak szybko się zwraca? Tak, subskrypcja zwraca się błyskawicznie dzięki dostępowi do funkcji Pro jak worktrees, nielimitowany Claude 3.5 Sonnet i tryb Composer. Nawet początkujący zyskują godziny tygodniowo, unikając manualnego debugowania i pisania boilerplate kodu. Koszt $20 to inwestycja w produktywność, a nie tylko wydatek na narzędzie.
### Jakie skróty klawiszowe w Cursor są najważniejsze dla zachowania płynności pracy i flow z AI? Absolutnym minimum jest `Cmd+E` do przełączania między kodem a agentem oraz `Shift+Tab` do szybkiej zmiany trybów (Ask/Agent/Plan). Warto też używać `Cmd+J` do togglowania terminala, co eliminuje odrywanie rąk od klawiatury. Te skróty pozwalają traktować AI jak naturalne rozszerzenie edytora.
### Jak skutecznie zarządzać limitem context window w Cursor aby uniknąć problemów z pamięcią agenta? Kluczem jest monitorowanie wskaźnika użycia i resetowanie czatu po przekroczeniu 60% pojemności, co zapobiega halucynacjom modelu. Należy też wyłączać nieużywane MCP w ustawieniach, ponieważ każdy aktywny dodatek konsumuje cenne tokeny. Precyzyjne używanie `@Files` zamiast całych folderów również oszczędza kontekst.
### Na czym polega praca z worktrees w Cursor i w jakich sytuacjach najlepiej ją stosować? Worktrees umożliwiają równoległe generowanie i testowanie kilku wariantów kodu (np. przez różne modele AI) na izolowanych branchach. Cursor automatycznie stawia środowiska dla każdej wersji, co pozwala w kilka minut porównać np. trzy podejścia do UI. To idealne narzędzie do eksperymentowania i podejmowania decyzji architektonicznych.
### Dlaczego warto dodawać własną dokumentację do indeksu Cursor zamiast polegać na ogólnej wiedzy modelu? Modele AI często mają nieaktualną wiedzę o najnowszych wersjach bibliotek (np. Tailwind v4), co prowadzi do generowania błędnego kodu. Dodanie URL dokumentacji w "Docs" sprawia, że agent korzysta z "single source of truth" i nie zmyśla nieistniejących funkcji. To eliminuje problem "cutoff date" i poprawia jakość generowanego kodu.
--- # 5 technik które zmienią sposób pracy z Claude Code Source: https://pawel.lipowczan.pl/blog/5-technik-pracy-z-claude-code Published: 2026-01-09 Istnieje bardzo duże prawdopodobieństwo, że zostawiasz większość potencjału swojego asystenta kodowania AI na stole. Gdy zaczynałem pracę z Claude Code przy budowie tego portfolio, robiłem dokładnie to samo - wpisywałem proste prompty, otrzymywałem kod, czasami działał, czasami nie. Reaktywne promptowanie. Bez systemu. Potem odkryłem, że najlepsi inżynierowie AI pracują zupełnie inaczej. Mają **system**. Wykorzystują metodologie, które sprawiają, że ich agenci kodowania stają się coraz potężniejsi z każdą iteracją. W tym artykule pokażę Ci 5 konkretnych technik, które całkowicie zmieniają sposób pracy z Claude Code. To nie są teoretyczne koncepcje - to praktyczne metody używane przez zespoły, które budują produkcyjne aplikacje z pomocą AI. A najlepsze jest to, że wszystkie te techniki są już spakowane w gotowy do użycia framework, który możesz wdrożyć w swoim projekcie jeszcze dziś. ## PRD-first development: Gwiazda polarna Twojego projektu Większość programistów po prostu nurkuje w kod. Otwierają Claude Code i zaczynają: "Dodaj przycisk logowania", "Zrób walidację formularza", "Napraw ten bug". Każda iteracja jest odseparowana od poprzedniej. Nie ma wizji całości. **PRD (Product Requirement Document)** w kontekście pracy z AI to coś znacznie prostszego niż korporacyjny dokument na 50 stron. To po prostu **markdown z pełnym zakresem projektu**. Pojedynczy plik, który staje się gwiazdą polarną dla każdej feature'ki, którą budujesz. ### Zalety PRD-first development Wcześniej korzystałem z asystenta, który generował PRD oraz rules dla agenta AI, ale to wymagało korzystania z dwóch różnych narzędzi. Musiałem przełączać się między kontekstami, synchronizować informacje ręcznie. To było frustrujące. Teraz? Wszystko w jednym miejscu. PRD żyje w moim repozytorium. Claude Code czyta go na początku każdej sesji. I nagle wszystko ma sens. ### Jak to wygląda w praktyce Dla nowych projektów (greenfield development), PRD zawiera: - **Target Users** - dla kogo budujesz - **Mission** - co ma robić ten produkt - **In Scope / Out of Scope** - co jest w MVP, a co na później - **Architecture** - high-level tech stack i struktura Dla istniejących projektów (brownfield), PRD dokumentuje: - **Co już mamy** - obecny stan systemu - **Co budujemy dalej** - kolejne feature'ki w pipeline - **Długoterminowa wizja** - gdzie zmierzamy Przykładowa minimalna struktura PRD: ```markdown # PRD: Habit Tracker Application ## Target Users Osoby chcące budować lepsze nawyki przez konsekwentne trackowanie ## Mission Prosta, elegancka aplikacja do śledzenia nawyków z wizualizacją postępów ## In Scope (MVP) - Tworzenie nawyków z nazwą i częstotliwością - Zaznaczanie wykonania nawyku dziennie - Kalendarz pokazujący historię (streak tracking) - Lokalne przechowywanie danych ## Out of Scope (v1) - Współdzielenie z innymi użytkownikami - Zaawansowane statystyki - Powiadomienia push - Integracje z innymi aplikacjami ## Architecture - Frontend: React + TypeScript - State: React Context - Storage: localStorage - Deploy: Vercel ``` ### Magiczne pytanie Gdy masz PRD, możesz zaczynać każdą sesję od pytania, które zmienia wszystko: > "Based on PRD, what should we build next?" Claude Code czyta PRD, rozumie gdzie jesteś, co już zbudowałeś, i sugeruje kolejny logiczny krok. Nie musisz pamiętać. Nie musisz tłumaczyć kontekstu od nowa. **PRD pamięta za Ciebie**. ### Kluczowe korzyści - **Single source of truth** - jedna definicja projektu dla całego zespołu (i dla AI) - **Naturalna dekompozycja** - łatwo wyodrębniasz kolejne feature'ki do zaimplementowania - **Context dla agenta** - Claude Code zawsze wie, nad czym pracujesz - **Spojrzenie z lotu ptaka** - każda feature łączy się z większą wizją PRD to fundament. Bez niego budujesz dom na piasku. Z nim - masz solidny fundament - każda linia kodu ma sens i kierunek. ## Modularność reguł: Lżejszy kontekst, mądrzejszy agent Widziałem to dziesiątki razy. Programista tworzy `CLAUDE.md` albo `agents.md`, wrzuca tam wszystkie możliwe reguły, wytyczne, konwencje. Po dwóch miesiącach plik ma 1500 linijek. Ladowane jest to **na początku każdej rozmowy**. Problem? **Przytłaczasz LLM nieistotnym kontekstem.** ### Problem długich reguł globalnych Gdy pracujesz nad frontend'em, nie potrzebujesz znać wzorców projektowania API. Gdy pracujesz nad bazą danych, nie musisz mieć w kontekście konwencji nazewnictwa komponentów React. Ale jeśli wszystko jest w jednym pliku globalnych reguł? Claude Code ładuje to wszystko. Za każdym razem. To marnowanie **okna kontekstu** - czegoś, co wielu programistów mocno niedocenia, a co jest absolutnie kluczowe dla jakości outputu agenta. ### Rozwiązanie: Architektura modularna Zamiast jednego monstrualnego pliku, dzielisz reguły na dwie kategorie: **1. Global rules (CLAUDE.md)** - maksymalnie lekkie, ~200 linijek - Tech stack projektu - Struktura folderów - Komendy do uruchomienia (npm run dev, npm test) - Strategia testowania (filozofia, nie szczegóły) - Standardy logowania **2. Reference folder** - szczegółowy kontekst ładowany tylko gdy potrzebny - `reference/api-design.md` - wzorce REST API, error handling, ładowane tylko przy pracy nad API - `reference/frontend-components.md` - component patterns, styling guidelines, tylko dla UI work - `reference/database-patterns.md` - schema design, migrations, query optimization, tylko dla DB work - `reference/testing-patterns.md` - szczegółowe przykłady testów, setup, mocking ### Jak to skonfigurować W swoim głównym `CLAUDE.md` dodajesz sekcję reference: ```markdown # CLAUDE.md ## Tech Stack - React 19 + TypeScript - Tailwind CSS 3 - Vite 7 ## Project Structure src/ ├── components/ ├── pages/ ├── utils/ └── data/ ## Reference Documentation When working on specific areas, consult these documents: - **API endpoints**: `.claude/reference/api-design.md` - **Frontend components**: `.claude/reference/frontend-components.md` - **Database operations**: `.claude/reference/database-patterns.md` - **Testing**: `.claude/reference/testing-patterns.md` ``` Claude Code jest wystarczająco inteligentny, żeby zrozumieć: "OK, pracuję teraz nad API endpoint'em, powinienem przeczytać api-design.md". I robi to automatycznie. ### Zalety modularności reguł Posiadanie tego wszystkiego w jednym miejscu (jak w claude-piv-skeleton, o którym zaraz opowiem) vs. rozproszenie po różnych narzędziach to **ogromna różnica**. Wszystko jest w repo. Wszystko jest wersjonowane. Wszystko ewoluuje razem z projektem. A najważniejsze? **Chronisz okno kontekstu** dla rzeczy, które naprawdę mają znaczenie podczas implementacji konkretnej feature'ki. ## Komendyfikacja: Skończ z powtarzaniem tych samych promptów Jeśli wysyłasz ten sam prompt do swojego agenta kodowania więcej niż dwa razy, to krzyczący sygnał: **"Zamień mnie w komendę!"** ### Co to są komendy Komendy to po prostu **pliki markdown, które definiują workflow**. Ładujesz je jako kontekst, a Claude Code wykonuje zdefiniowany proces krok po kroku. To jak makra dla promptów. Nie wymaga to żadnych nowych narzędzi. To po prostu lepszy sposób organizacji pracy. ### Co warto skomendyfikować Praktycznie wszystko, co robisz regularnie: **Core workflow:** - `/prime` - Załaduj kontekst projektu na początku sesji - `/plan-feature` - Stwórz strukturalny plan implementacji feature'ki - `/execute` - Zaimplementuj według planu - `/validate` - Uruchom testy i zweryfikuj działanie **Git operations:** - `/commit` - Stwórz sensowny commit message na podstawie zmian - `/review-pr` - Przeanalizuj pull request przed merge **Maintenance:** - `/update-docs` - Zaktualizuj dokumentację po zmianach w kodzie - `/refactor` - Zrefaktoruj zgodnie z najlepszymi praktykami projektu ### Przykład: Komenda /prime ```markdown # Command: /prime Load codebase context to prepare for feature development work. ## Objective Ensure Claude Code understands current project state before starting any development. ## Steps 1. Read PRD from `docs/PRD.md` to understand project scope and vision 2. Read architecture documentation from `docs/ARCHITECTURE.md` 3. Scan recent commits: `git log -10 --oneline` to see latest changes 4. List current todos from `docs/TODO.md` to understand priorities 5. Check git status to see any uncommitted work 6. Confirm context successfully loaded with summary of project state ## Expected Output Brief summary including: - Project name and current phase - Last 3 features implemented - Current priorities from TODO - Any blockers or issues noted ``` Zamiast każdorazowo wypisywać: "Przeczytaj PRD, potem architecture, potem sprawdź co się ostatnio zmieniło..." - po prostu wywołujesz `/prime` i gotowe. ### Kluczowa obserwacja Ponieważ asystent to praktycznie tylko prompt, można ten prompt wykorzystać jako command. To piękna właściwość markdown'owych komend - są przenośne, czytelne dla człowieka, i możesz je iteracyjnie ulepszać. ### Typowy workflow z komendami Tak wygląda mój dzienny cykl pracy z Claude Code: ```text /prime ↓ "Based on PRD, what should we build next?" ↓ /plan-feature "Add user authentication" ↓ [Context Reset - nowa konwersacja] ↓ /execute plan-auth.md ↓ /validate ↓ /commit ``` ### Korzyści komendyfikacji - **Konsystencja** - ten sam proces za każdym razem, zero pominięć kroków - **Zero zapominania** - nie musisz pamiętać sekwencji, komenda pamięta - **Onboarding** - nowy programista w zespole dostaje gotowe komendy i od razu działa produktywnie - **Ciągłe ulepszanie** - komendy ewoluują, stają się lepsze z czasem Oszczędzasz dosłownie **tysiące naciśnięć klawiszy** rocznie. A co ważniejsze - oszczędzasz energię mentalną na rzeczy, które naprawdę wymagają myślenia. ## Reset kontekstu: Najważniejszy krok którego nie robisz To brzmi wbrew intuicji: **Zawsze resetuj konwersację między planowaniem a wykonaniem.** Większość programistów robi to źle. Planują feature'kę w długiej rozmowie z Claude Code - czytają pliki, dyskutują o architekturze, eksplorują różne podejścia. A potem, w tej samej rozmowie, od razu zaczynają implementację. Problem? **Okno kontekstu jest zaśmiecone** całym procesem eksploracji. ### Zalety resetu kontekstu Podczas planowania ładujesz TONY kontekstu: - Czytasz wiele plików z różnych części projektu - Eksplorujesz istniejącą architekturę - Dyskutujesz o różnych podejściach - Przeglądasz podobne implementacje w codebase Ale podczas egzekucji chcesz: - **Maksymalną przestrzeń do rozumowania** dla LLM - **Miejsce na self-validation** - agent powinien móc sprawdzać swoje własne rozwiązania - **Czysty mental model** - tylko to, co jest potrzebne do implementacji tej konkretnej feature'ki ### Właściwy workflow ```text [Planning Session] ↓ /prime - Załaduj kontekst codebase ↓ Rozmowa: "Based on PRD, let's plan authentication feature" ↓ /plan-feature - Output: structured plan jako markdown document ↓ [NOWA KONWERSACJA] ← To jest kluczowe! ↓ /execute plan-auth.md - JEDYNY kontekst to plan ↓ Implementacja z pełną przestrzenią do reasoning ``` ### Co zawiera ten plan document Plan musi być self-contained - zawierać WSZYSTKO, co potrzebne do implementacji: ```markdown # Plan: User Authentication Feature ## Feature Description Implement JWT-based authentication with email/password login. ## User Story As a user, I want to securely log in to access personalized content. ## Context to Reference - `src/utils/api.js` - API utility functions - `src/context/AuthContext.jsx` - existing auth context (modify) - `docs/reference/api-design.md` - API patterns ## Technical Approach 1. Backend: Create /api/auth/login and /api/auth/register endpoints 2. Frontend: Login form component with validation 3. State: Store JWT token in AuthContext 4. Protected routes: Add authentication middleware ## Task-by-Task Breakdown ### Task 1: Backend Auth Endpoints - Create `src/api/auth.js` with login and register functions - Implement JWT token generation - Add password hashing with bcrypt - Error handling for invalid credentials ### Task 2: Login Form Component - Create `src/components/LoginForm.jsx` - Form validation using Formik - Connect to auth API - Handle loading and error states ### Task 3: Auth Context Updates - Modify `src/context/AuthContext.jsx` - Add login/logout/register methods - Persist token to localStorage - Auto-refresh token logic ### Task 4: Protected Routes - Create ProtectedRoute component - Redirect to /login if not authenticated - Update router configuration ## Testing Requirements - Unit tests for auth API functions - Integration tests for login flow - E2E test: complete registration and login - Error handling tests: invalid credentials, expired token ## Success Criteria - User can register with email/password - User can login and access protected pages - Token persists across page refreshes - Logout clears token and redirects to home ``` ### Zalety planu self-contained LLM ma **maksymalną liczbę tokenów** dla reasoning podczas krytycznej fazy kodowania. Nie ma kontaminacji kontekstem eksploracyjnym. A co najważniejsze - **zmusza Cię to do tworzenia kompletnych, self-contained planów**. To dyscyplina. Ale dyscyplina, która czyni Twoje sesje kodowania z AI nieporównywalnie bardziej efektywnymi. Próbowałem obydwoma sposobami. Różnica jest gigantyczna. Reset kontekstu to technika, której uczyłem się najdłużej, ale która dała największy boost w jakości outputu. ## Ewolucja systemu: Każdy bug to lekcja dla agenta To najważniejsza technika ze wszystkich. I najczęściej pomijana. Dla mnie do niedawna nieznana. Być może dla Ciebie także, albo nie zdajesz sobie sprawy z jej znaczenia. Tradycyjne podejście wygląda tak: 1. Agent AI popełnia błąd 2. Naprawiasz go ręcznie 3. Idziesz dalej 4. **Ten sam błąd powtarza się za tydzień** Podejście ewolucyjne: 1. Agent AI popełnia błąd 2. Analizujesz: **Co w systemie pozwoliło na ten błąd?** 3. Aktualizujesz rules/commands/process 4. **Ta klasa bugów jest wyeliminowana na zawsze** ### Shift w mindset Nie naprawiaj buga. **Napraw system, który pozwolił na buga.** Twój agent AI to nie jest statyczne narzędzie. To **ewoluujący system**, który może stawać się coraz potężniejszy z każdą iteracją. Ale tylko jeśli Ty aktywnie go ulepszasz. ### Przykłady ewolucji systemu **Scenario 1: Złe style importów** ```text Bug: Agent używa require() zamiast import w projekcie ES6 Analiza: Brak jasnej reguły o module system Fix: Dodaj do CLAUDE.md: "Always use ES6 import/export syntax. Never use require() or module.exports." Rezultat: Nigdy więcej problem ze stylami importów ``` **Scenario 2: Zapomina uruchamiać testy** ```text Bug: Agent implementuje feature bez testów Analiza: Brak kroku "testing" w workflow Fix: Zaktualizuj template /execute command: ## Testing Phase 1. Write tests first (TDD when appropriate) 2. Run full test suite: npm test 3. Ensure all tests pass before completion 4. Add test coverage report to plan Rezultat: Testy stają się automatyczną częścią workflow ``` **Scenario 3: Nie rozumie auth flow** ```text Bug: Agent implementuje niepoprawną autentykację Analiza: Brak dokumentacji authentication flow Fix: 1. Stwórz reference/authentication.md z flow diagram 2. Dodaj do CLAUDE.md: "When working on authentication, read reference/authentication.md" Rezultat: Implementacje auth konsekwentnie poprawne ``` ### Workflow refleksji Po zakończeniu każdej feature'ki, zamiast od razu lecieć do następnej: ```markdown "Hey Claude, zauważyłem że XYZ nie działało poprawnie i musiałem to naprawić. Przeanalizujmy: 1. Przeczytaj komendy, których użyliśmy 2. Przeczytaj obecne rules 3. Zidentyfikuj co możemy ulepszyć, żeby to się nie powtórzyło 4. Zasugeruj konkretne zmiany w rules/commands" ``` Claude Code przeanalizuje session, znajdzie luki w procesie, i zaproponuje poprawki. Czasami to nowa reguła. Czasami dodatkowy krok w komendzie. Czasami nowy reference document. ### Zalety ewolucji systemu - **Twój agent staje się mądrzejszy z czasem** - reliability rośnie z każdą iteracją - **Budujesz institutional knowledge** - cały zespół korzysta z nauki na błędach - **Transformation z reactive do proactive** - zamiast gasić pożary, zapobiegasz im - **Compound effect** - po 3 miesiącach masz system, który robi mniej błędów niż junior developer To więcej niż technika. To **mindset**. Traktuj system swojego agenta AI jak kod produkcyjny, który wymaga ciągłego ulepszania. Z własnego doświadczenia wiem, że największym błędem jest ignorowanie wzorców w błędach agenta. Pierwszy raz to przypadek. Drugi raz to sygnał. Trzeci raz to Twoja wina, że nie poprawiłeś systemu. ## Claude PIV Skeleton: Wszystko w jednym miejscu Teraz najlepsza część. Wszystkie te techniki, o których mówię - PRD-first development, modularność reguł, komendyfikacja workflow, reset kontekstu, ewolucja systemu - są już zaimplementowane w gotowym do użycia framework. Nazywa się **Claude PIV Skeleton** i jest dostępny na GitHub: [https://github.com/plipowczan/claude-piv-skeleton](https://github.com/plipowczan/claude-piv-skeleton) (fork od [galando](https://github.com/galando/claude-piv-skeleton)) ### Problem, który rozwiązuje Wcześniej korzystałem z asystenta, który generował PRD oraz rules dla agenta AI, ale to wymagało korzystania z dwóch różnych narzędzi. Musiałem kopiować output między aplikacjami, synchronizować manualne, tracić kontekst. Posiadanie tego wszystkiego w jednym miejscu jest **ogromną zaletą**. A ponieważ asystent to praktycznie tylko prompt, można ten prompt wykorzystać jako command - i dokładnie to robi PIV Skeleton. ### Czym jest PIV Methodology PIV to akronim od **Prime-Implement-Validate** - metodologia stworzona przez [Cole Medin](https://github.com/coleam00) specjalnie dla development z asystentami AI: - **Prime**: Załaduj i zrozum kontekst codebase - **Implement**: Zaplanuj feature'ki i wykonaj implementację - **Validate**: Automatycznie testuj i weryfikuj To nie abstrakcja. To konkretny workflow, który Cole udowodnił w produkcyjnych projektach jak [woningscoutje.nl](https://woningscoutje.nl). ### Co dostajemy w claude-piv-skeleton Repository implementuje wszystkie 5 technik jako ready-to-use framework: **Universal methodology** - działa z każdym tech stackiem (Spring Boot, Node.js, React, Python FastAPI, etc.) **Modular rules system** - path-based loading reguł w zależności od tego, nad czym pracujesz **Pre-built commands** - kompletny zestaw komend dla całego PIV workflow **Technology templates** - gotowe konfiguracje dla popularnych stacków ### Struktura repozytorium ```text .claude/ ├── CLAUDE.md # Lightweight global rules ├── PIV-METHODOLOGY.md # Pełna dokumentacja metodologii ├── commands/ # Wszystkie workflows jako komendy │ ├── piv_loop/ # Core PIV workflow │ │ ├── prime.md # Prime phase command │ │ ├── plan-feature.md # Planning command │ │ └── execute.md # Execution command │ ├── validation/ # Validation & testing │ │ ├── validate.md # Full validation pipeline │ │ ├── code-review.md # Technical review │ │ └── system-review.md # Process improvement │ └── bug_fix/ # Bug fix workflow │ ├── rca.md # Root cause analysis │ └── implement-fix.md # Fix implementation ├── rules/ # Modular rules by technology │ ├── 00-general.md # Universal principles │ ├── 10-git.md # Git workflow │ ├── 20-testing.md # Testing philosophy │ └── backend/ # Backend-specific rules └── reference/ # Best practices loaded on-demand └── patterns/ # Design patterns reference ``` ### Workflow, który umożliwia Kompletny cykl development z PIV Skeleton: ```bash # 1. Prime workspace "Run /piv_loop:prime to load the project context" # 2. Plan feature "Use /piv_loop:plan-feature to create a plan for adding user authentication" # 3. Execute (automatic context reset!) "Use /piv_loop:execute to implement the plan" # 4. Validation runs automatically # No manual step needed - testing happens in the workflow # 5. Bug fix with system evolution built in "Run /bug_fix:rca for issue #123" "Use /bug_fix:implement-fix to implement the fix" ``` ### Zalety PIV Skeleton - **Nie musisz budować od zera** - gotowa struktura, przetestowana w produkcji - **Battle-tested patterns** - workflow wypracowany przez najlepszych AI engineers - **Community-driven** - wkład wielu programistów, ciągłe ulepszenia - **Extensible** - łatwo dostosować do swojego tech stacku i potrzeb PIV Skeleton to nie tylko kod. To **system myślenia** o pracy z AI agents, zapakowany w reusable framework. ## Jak zacząć: Pierwsze kroki Świetnie, znasz już 5 technik i wiesz, że istnieje gotowy framework. Ale jak to wszystko wdrożyć w praktyce? ### Ścieżka 1: Użyj PIV Skeleton (Recommended) Jeśli zaczynasz nowy projekt albo możesz przenieść istniejący: ```bash # Sklonuj repozytorium git clone https://github.com/plipowczan/claude-piv-skeleton.git my-project cd my-project # Usuń historię git, żeby zacząć od czystej kartki rm -rf .git git init # Zainstaluj swój tech stack # (Postępuj według technology-specific guides w technologies/ directory) # Rozpocznij pierwszą feature'kę # Otwórz Claude Code i: ``` 1. `"Run /piv_loop:prime to load project context"` 2. `"Based on PRD, what should we build first?"` 3. `"Use /piv_loop:plan-feature to plan it"` 4. `"Use /piv_loop:execute to implement"` I to wszystko. Masz gotowy system z pierwszą feature'ką w produkcji. ### Ścieżka 2: Implementuj Inkrementalnie Jeśli masz istniejący projekt i nie chcesz wszystkiego przenosić naraz, wdrażaj techniki krok po kroku: **Tydzień 1: Stwórz PRD** - Udokumentuj obecny stan projektu - Zdefiniuj kolejne feature'ki do zbudowania - Uczyń PRD swoją gwiazdą polarną **Tydzień 2: Stwórz komendę Prime** - Jaki kontekst powinien być zawsze ładowany? - Stwórz `/prime` command w `.claude/commands/` - Używaj na początku każdej sesji **Tydzień 3: Modularyzuj Rules** - Podziel swój CLAUDE.md na global + reference - Przenieś task-specific rules do `reference/` - Dodaj sekcję reference w global rules **Tydzień 4: Dodaj Feature Workflow** - Stwórz `/plan-feature` command - Ćwicz context reset między planem a wykonaniem - Stwórz `/execute` command **Tydzień 5: System Evolution** - Po każdym bugu, zrób refleksję - Aktualizuj rules/commands na podstawie wniosków - Trackuj improvement w CHANGELOG ### Kluczowe czynniki sukcesu 1. **Start small** - Nie próbuj wdrożyć wszystkich 5 technik naraz. Zacznij od PRD, potem dodaj prime command, etc. 2. **Dokumentuj na bieżąco** - Zapisuj co działa, co nie. Twoje notatki staną się częścią ewolucji systemu. 3. **Iteruj na komendach** - Twoje workflows będą się ulepszać. To normalne. Po miesiącu Twój `/prime` będzie lepszy niż na początku. 4. **Bądź konsekwentny** - Używaj systemu za każdym razem. Nie wracaj do starych nawyków "szybkiego fix'a" bez procesu. 5. **Udostępnij zespołowi** - Jeśli pracujesz w zespole, upewnij się że wszyscy używają tych samych komend i procesów. Mnożysz korzyści. ### Twój pierwszy feature z PIV Konkretny przykład - załóżmy, że budujesz habit tracker i chcesz dodać streak tracking: ```text Session 1 - Planning: → /prime → "Based on PRD, let's plan streak tracking feature" → /plan-feature "Streak tracking - show consecutive days" → Plan zapisany: .claude/agents/plans/streak-tracking.md [Restart konwersacji] Session 2 - Execution: → /execute .claude/agents/plans/streak-tracking.md → [Implementation happens with tests] → /validate → All tests pass ✓ Session 3 - Commit: → /commit → "feat: Add streak tracking with visual indicators" ``` Po godzinie masz feature'kę w produkcji. Z testami. Z poprawnym commit message. Z wszystkim. ## Kluczowe wnioski 1. **PRD-first development** zapewnia spójność i kierunek dla wszystkich iteracji z agentem AI. To gwiazda polarna, która sprawia, że każda feature ma sens w kontekście całości. 2. **Modularyzacja reguł** chroni okno kontekstu i ładuje tylko potrzebną wiedzę. Przestań marnować tokeny na nieistotny kontekst - ładuj to, co ważne, wtedy gdy ważne. 3. **Komendyfikacja workflow** oszczędza tysiące naciśnięć klawiszy i zapewnia konsystencję. Jeśli robisz coś więcej niż dwa razy, to powinno być komendą. 4. **Reset kontekstu** między planowaniem a wykonaniem daje agentowi maksymalną przestrzeń do rozumowania. Counterintuitive, ale to jedna z najbardziej impactowych technik. 5. **Ewolucja systemu** przekształca każdy bug w lekcję, która sprawia że agent staje się mądrzejszy. Nie naprawiaj buga - napraw system, który na niego pozwolił. 6. **PIV Skeleton** oferuje gotową implementację wszystkich technik w jednym miejscu. Nie musisz budować od zera - możesz zacząć już dziś. 7. Najważniejsze: **systematyczne podejście** vs. reaktywne promptowanie to różnica między wykorzystaniem 20% a 80% potencjału Claude Code. ### Moje doświadczenie Gdy zaczynałem budować to portfolio z pomocą AI, nie miałem systemu. Proste prompty, ad-hoc fixes, zero procesów. Potem odkryłem te techniki. I wszystko się zmieniło. Teraz mój workflow z Claude Code jest przewidywalny. Efektywny. I co najważniejsze - **agent staje się lepszy z każdą sesją**, zamiast popełniać te same błędy w kółko. To transformacja, którą możesz mieć w swoim zespole. Wymaga to zmiany mindset z "AI to szybszy Google" na "AI to evolving development partner". Ale jeśli to zrobisz? Różnica będzie gigantyczna. ## FAQ
### Jak zacząć z PRD-first development jeśli nigdy wcześniej nie tworzyłem dokumentów projektowych? Zacznij od minimalnego PRD z czterema sekcjami: Target Users (dla kogo), Mission (co robi), In Scope (MVP features) i Out of Scope (co na później). Nie potrzebujesz 50-stronicowego dokumentu - wystarczy prosty markdown z kluczowymi decyzjami. PRD dla małego projektu może mieć dosłownie 20-30 linijek i już daje ogromną wartość jako single source of truth dla AI.
### Ile dokładnie reguł powinienem mieć w głównym pliku CLAUDE.md żeby nie przytłoczyć kontekstu LLM? Maksymalnie 200 linijek w głównym CLAUDE.md - tech stack, struktura projektu, podstawowe komendy i linki do reference docs. Wszystkie szczegółowe wzorce (API design, component patterns, testing) przenoś do osobnych plików w folderze reference. Claude Code automatycznie załaduje je tylko wtedy gdy pracujesz nad danym obszarem, oszczędzając cenne miejsce w oknie kontekstu.
### Czy komendyfikacja workflow działa tylko z Claude Code czy mogę używać tych samych komend z ChatGPT lub innymi LLM? Komendy to zwykłe pliki markdown z instrukcjami workflow, więc działają z dowolnym LLM (ChatGPT, Claude, Cursor, Windsurf). Jedyna różnica to sposób ładowania - w Claude Code to slash commands, w ChatGPT kopiujesz zawartość jako prompt. Sama metodologia i struktura komend jest uniwersalna i przenośna między narzędziami.
### Dlaczego muszę resetować kontekst między planowaniem a wykonaniem zamiast zrobić wszystko w jednej sesji? Podczas planowania ładujesz TONY eksploracyjnego kontekstu (czytasz wiele plików, dyskutujesz o różnych podejściach), który zaśmieca okno kontekstu LLM. Reset daje agentowi czysty slate z maksymalną przestrzenią do rozumowania i self-validation podczas implementacji. To counterintuitive, ale empirycznie daje znacznie lepsze rezultaty - agent ma miejsce na quality checks zamiast walczyć z przeładowanym kontekstem.
### Czym dokładnie różni się claude-piv-skeleton od zwykłego używania Claude Code bez żadnego systemu? PIV Skeleton to gotowy framework z predefiniowanymi komendami (/prime, /plan-feature, /execute, /validate), strukturą folderów (.claude/commands, .claude/agents), szablonami dokumentów (PRD, rules) i procesem ewolucji systemu. Zamiast wymyślać workflow od zera, dostajesz sprawdzony system używany przez teams w produkcji - po prostu fork'ujesz repo i masz gotowe best practices. To jak różnica między pisaniem własnego framework'a a użyciem Next.js.
### Ile czasu realistycznie zajmuje wdrożenie tych 5 technik w istniejącym projekcie który już trwa kilka miesięcy? Start small - dzień 1: stwórz minimalny PRD (1-2h), dzień 2: lekki CLAUDE.md + jedna komenda /prime (1-2h), tydzień 1: dodaj /plan-feature i /execute (2-3h łącznie). Nie implementuj wszystkiego naraz. Po 2 tygodniach pracy z systemem zobaczysz naturalne miejsca do dodania kolejnych komend i reguł. Brownfield projects wymagają około 5-8 godzin total setup, ale ROI widzisz już po pierwszej sesji z nowym workflow.
---

Chcesz wdrożyć AI w swoim zespole?

Pomogę Ci zbudować system pracy z agentami AI, który zwiększy produktywność Twojego zespołu. Od strategii przez implementację po szkolenia.

Umów bezpłatną konsultację
## Przydatne zasoby - **[claude-piv-skeleton](https://github.com/plipowczan/claude-piv-skeleton)** - Gotowy framework implementujący PIV methodology - **[habit-tracker](https://github.com/coleam00/habit-tracker)** - Oryginalny projekt demo PIV od Cole Medin - **[context-engineering-intro](https://github.com/coleam00/context-engineering-intro)** - Wprowadzenie do context engineering od Cole Medin - **[Claude Code Documentation](https://claude.ai/code)** - Oficjalna dokumentacja Claude Code --- # 2026: Rok, w którym AI przeszła z laboratoriów do hal produkcyjnych Source: https://pawel.lipowczan.pl/blog/trendy-ai-2026-od-eksperymentow-do-operacjonalizacji Published: 2026-01-01 Siedzimy w pierwszym dniu 2026 roku. Jeśli jesteś liderem technologicznym, decision makerem w firmie lub po prostu kimś, kto próbuje nadążyć za rewolucją AI, to prawdopodobnie czujesz mieszankę ekscytacji i niepewności. I dobrze. Bo rok 2026 to moment, w którym AI przestaje być "fascynującą technologią przyszłości", a staje się fundamentem operacyjnym - narzędziem, które albo zintegrujemy z naszymi procesami biznesowymi, albo zostaniemy w tyle. Po latach eksperymentów, pilotaży i prezentacji "wow effect" przychodzi czas na weryfikację. Analitycy Gartnera nazywają to "cyklem superinteligencji", Forrester mówi o "otrzeźwieniu" (the reckoning), a Deloitte o "infrastrukturalnym rozrachunku". Ja nazywam to po prostu: **końcem turystyki AI i początkiem prawdziwej pracy**. ## Przejście od chatbotów do agentów: AI, która działa zamiast tylko odpowiadać Największa zmiana, którą obserwuję w 2026 roku, to przesunięcie od modeli generatywnych jako "mądrych asystentów" w stronę **agentowej AI (Agentic AI)** - systemów, które nie tylko odpowiadają na pytania, ale samodzielnie planują, podejmują decyzje i wykonują działania. ### Co to właściwie oznacza w praktyce? Wyobraź sobie, że prosisz AI: "Przeprowadź audyt zgodności RODO dla nowego procesu marketingowego". Wcześniejsze modele (GPT-4, wczesne wersje Claude) dawałyby Ci listę kroków do wykonania. **Agenci w 2026 roku wykonują te kroki samodzielnie**: 1. Agent "badacz" analizuje dokumentację procesu 2. Agent "prawnik" sprawdza zgodność z przepisami RODO 3. Agent "kodyfikator" generuje checklisty i raporty 4. Agent "krytyk" weryfikuje wnioski i eskaluje wątpliwości do człowieka To już nie jest science fiction. IDC prognozuje, że do 2029 roku systemy agentyczne będą odpowiadać za blisko 50% wszystkich wydatków na AI. Do końca 2026 roku, 80% aplikacji enterprise będzie zawierało wbudowanych agentów AI. ### Orkiestracja wieloagentowa: zespoły AI w akcji Kluczowa zmiana architektoniczna to przejście od pojedynczych, monolitycznych modeli do **systemów wieloagentowych** (Multi-agent Systems). Zamiast jednego "super-mózgu", mamy zespół wyspecjalizowanych agentów, którzy współpracują ze sobą przez protokoły takie jak Model Context Protocol (MCP) czy Agent-to-Agent (A2A). W praktyce widziałem już wdrożenia w: - **Obsłudze klienta**: Systemy agencyjne w bankowości zapewniają wsparcie 24/7, automatycznie koordynując działania między działami - **Marketingu**: AI generuje briefy kampanii, segmentuje odbiorców, testuje warianty i optymalizuje budżety bez udziału człowieka - **Rozwoju oprogramowania**: Agenci automatyzują testy, zgłaszają bugi, generują dokumentację i nawet proponują poprawki kodu Ale uwaga - **sukces Agentic AI zależy od "ograniczonej autonomii"**. Agenci muszą działać w ściśle zdefiniowanych ramach bezpieczeństwa, z mechanizmami eskalacji do człowieka w przypadku anomalii. To odpowiedź na problemy z halucynacjami wcześniejszych modeli. ## Modele rozumowania: AI, która "myśli" zanim odpowie Jeśli Agentic AI to rewolucja w działaniu, to **modele rozumowania (Reasoning Models)** to rewolucja w myśleniu. W 2026 roku AI nie generuje już odpowiedzi natychmiastowo - zamiast tego, poświęca dodatkowy czas obliczeniowy na "namysł". ### System 2: wolne, analityczne myślenie AI Termin "System 2" pochodzi z psychologii kognitywnej i odnosi się do procesów myślowych, które są wolne, analityczne i logiczne (w przeciwieństwie do szybkiego, intuicyjnego Systemu 1). Modele nowej generacji - następcy OpenAI o1/o3, Google Gemini w zaawansowanych wersjach - wykorzystują **inference-time compute**: przeprowadzają wewnętrzne symulacje, weryfikują hipotezy i planują kroki rozwiązania przed wygenerowaniem ostatecznej odpowiedzi. Efekty? AI osiąga poziom ekspercki w: - Rozwiązywaniu problemów matematycznych na poziomie 93%+ dokładności (GPQA Diamond) - Zadaniach wieloetapowych wymagających logicznego rozumowania - Weryfikacji poprawności założeń i wykrywaniu błędów w argumentacji Microsoft określa to jako przejście od "AI jako narzędzia" do "AI jako partnera", który nie tylko wykonuje polecenia, ale także wnosi wkład merytoryczny. ### Koszt inteligencji Ale jest haczyk. Gemini 3 z włączonym rozumowaniem zużył 160 milionów tokenów tam, gdzie bez rozumowania wystarczyło 7,4 miliona. To ilustruje fundamentalny **trade-off między szybkością a inteligencją**, który organizacje muszą aktywnie zarządzać w 2026 roku poprzez "budżety rozumowania" dostosowane do konkretnych zadań. ## Specjalizacja modeli: koniec ery "jednego modelu do wszystkiego" Jednym z najważniejszych trendów 2026 roku jest masowe przejście od wielkich, uniwersalnych modeli językowych (LLM) do **modeli specyficznych dla domeny (DSLM - Domain-Specific Language Models)**. ### Dlaczego specjalizacja wygrywa? W branżach regulowanych - medycynie, finansach, prawie - **dokładność jest ważniejsza niż uniwersalność**. Modele DSLM oferują: - **Wyższą precyzję**: Med-PaLM osiąga 95% dokładności w diagnostyce medycznej, FinGPT redukuje wykrywanie fraudów o 30%, JurisGPT analizuje kontrakty o 25-30% dokładniej niż LLM ogólnego przeznaczenia - **Niższe koszty operacyjne**: Mniejsza liczba parametrów oznacza redukcję kosztów inferencji nawet o 45% - **Wbudowaną zgodność z regulacjami**: Modele trenowane na dedykowanych zbiorach danych zawierają mechanizmy compliance "z pudełka" Gartner prognozuje, że do końca 2026 roku ponad 50% modeli GenAI wykorzystywanych przez przedsiębiorstwa będzie specyficznych dla danej dziedziny. W sektorach regulowanych ten odsetek sięga 80-90%. ### Architektura hybrydowa jako standard W praktyce nie widzę totalnego zastąpienia LLM przez DSLM. Zamiast tego obserwuję **architekturę hybrydową**: modele ogólnego przeznaczenia do szerokich zadań + moduły domenowe do specjalistycznych funkcji. Cloud providers (AWS, Azure, Google Cloud) już oferują dedykowane platformy: Healthcare AI, Financial Services AI, Manufacturing AI - każda pre-trenowana na kurowanych zbiorach danych z wbudowanymi frameworkami zgodności. ## Infrastruktura: od chmury do brzegu sieci Jeśli modele to mózg AI, to infrastruktura to jej ciało. I w 2026 roku to ciało przechodzi dramatyczną transformację. ### Ekonomia wnioskowania vs. trenowania Kluczowa zmiana: o ile wcześniej większość mocy obliczeniowej pochłaniał **trening modeli**, o tyle teraz dominującym obciążeniem staje się **wnioskowanie (inference)**. Deloitte przewiduje, że wnioskowanie będzie odpowiadać za dwie trzecie całkowitego zapotrzebowania na moc obliczeniową AI. To napędza boom na: - **Układy ASIC** (Application-Specific Integrated Circuits): AWS Trainium/Inferentia, Google TPU v6, Microsoft Maia - zoptymalizowane pod konkretne architektury modeli, oferujące lepszy stosunek wydajności do zużycia energii - **Trójwarstwową architekturę hybrydową**: 1. **Chmura publiczna**: elastyczność dla treningu i eksperymentów 2. **On-premises**: stabilność dla krytycznego wnioskowania i zgodność z suwerennością danych 3. **Edge/brzeg sieci**: ultra-niska latencja, prywatność, odporność na awarie centralnych usług ### Edge AI i TinyML: inteligencja wszędzie Edge computing i technologie takie jak TinyML (Tiny Machine Learning) redefiniują sposób przetwarzania danych w 2026 roku. Modele ML uruchamiane na **mikrokontrolerach i urządzeniach IoT** umożliwiają: - **Analizę w czasie rzeczywistym** bez wysyłania danych do chmury - **Niskie zużycie energii** (architektury neuromorficzne Intel Loihi 2 zużywają rzędy wielkości mniej energii) - **Ochronę prywatności** (dane medyczne, finansowe pozostają lokalnie) Praktyczne zastosowania już działają: - Smart agriculture: edge-deployed modele monitorują uprawy w czasie rzeczywistym - Predictive maintenance: wykrywanie anomalii sprzętu bezpośrednio na czujnikach - Wearables: analiza zdrowia bez ciągłego wysyłania danych do chmury ### AI PC i małe modele językowe W 2026 roku Gartner przewiduje, że 55% wszystkich nowych komputerów to **AI PC** wyposażone w dedykowane układy NPU (Neural Processing Unit) o wydajności przekraczającej 40-50 TOPS. Równolegle, rynek mobilny przeżywa renesans dzięki **Małym Modelom Językowym (SLM)** - Google Gemini Nano, Apple Intelligence - liczącym od 1 do 7 miliardów parametrów, zoptymalizowanym do działania na procesorach mobilnych. Ponad połowa nowych smartfonów w 2026 roku ma natywne wsparcie GenAI, umożliwiając funkcje RAG (Retrieval-Augmented Generation) bezpośrednio na telefonie. To tworzy nową jakość "osobistej AI", która zna kontekst użytkownika, ale nie dzieli się nim z korporacjami. ## Physical AI: od robotów demonstracyjnych do produkcyjnych 2026 rok to moment, w którym **Physical AI** - sztuczna inteligencja posiadająca ciało - wkracza do hal produkcyjnych i magazynów na skalę komercyjną. ### Roboty humanoidalne: Tesla Optimus, Figure AI, Digit - **Tesla Optimus**: Elon Musk celuje w 2026 jako moment rozpoczęcia seryjnej produkcji i dostępności dla klientów zewnętrznych. Roboty przejmują proste, powtarzalne i niebezpieczne zadania - **Figure AI + BMW**: Partnerstwo osiąga dojrzałość - roboty Figure 02 pracują autonomicznie na liniach montażowych BMW, wykonując zadania manipulacyjne wymagające precyzji - **Agility Robotics (Digit)**: Robot znany z pracy w centrach logistycznych Amazon i GXO osiąga skalowalność operacyjną dzięki autonomicznemu dokowaniu i integracji z systemami WMS Kluczem jest "uniwersalny mózg robota" - model AI, który pozwala maszynie uczyć się nowych zadań poprzez obserwację, a nie imperatywne programowanie. ### Software-Defined Factory W przemyśle następuje integracja fizycznej automatyki z cyfrową inteligencją. Koncepcja **Software-Defined Factory** zakłada, że funkcjonalność linii produkcyjnej jest definiowana przez oprogramowanie. IDC prognozuje, że do 2029 roku 30% fabryk będzie zarządzanych przez otwarte platformy automatyki. AI przestaje być dodatkiem do predykcyjnej konserwacji, a staje się **systemem autonomicznym zarządzającym harmonogramowaniem produkcji** - ponad 40% producentów zmodernizuje swoje systemy planowania o moduły AI reagujące dynamicznie na zakłócenia w łańcuchu dostaw. ## EU AI Act: sierpień 2026 - godzina zero dla compliance Dla firm operujących w Europie najważniejszą datą kalendarza jest **2 sierpnia 2026 roku** - termin pełnej implementacji przepisów EU AI Act dotyczących systemów AI wysokiego ryzyka. ### Co to oznacza w praktyce? Od tego dnia firmy muszą mieć wdrożone: 1. **Systemy zarządzania ryzykiem AI** 2. **Mechanizmy nadzoru ludzkiego** (Human-in-the-loop) 3. **Gwarancję jakości danych treningowych** 4. **Pełną dokumentację techniczną i logi systemowe** 5. **Procedury raportowania incydentów** Brak zgodności? Kary sięgają **35 mln euro lub 7% globalnego obrotu** - to stawia compliance AI na równi z RODO jako priorytet zarządczy. ### Polska implementacja Projekt ustawy wdrażającej AI Act przewiduje powołanie Komisji Rozwoju i Bezpieczeństwa Sztucznej Inteligencji oraz uruchomienie pierwszej piaskowicy regulacyjnej do sierpnia 2026. To szansa dla polskich firm na bezpieczne testowanie rozwiązań AI w kontrolowanych warunkach. ## Cyberbezpieczeństwo: od reaktywnego do prewencyjnego Krajobraz zagrożeń w 2026 roku jest zdominowany przez ataki wspomagane AI. **Deepfake'i audio i wideo** są używane do omijania biometrii i zaawansowanego phishingu (fałszywe wideokonferencje z zarządem). ### Platformy bezpieczeństwa AI Gartner prognozuje, że do 2028 roku ponad połowa firm wdroży **platformy bezpieczeństwa AI (AI Security Platforms)**, które: - Centralizują widoczność wszystkich systemów AI w organizacji - Egzekwują polityki użycia AI (AI Usage Control) - Chronią przed zagrożeniami specyficznymi dla AI: prompt injection, wyciek danych, manipulacja agentów - Monitorują działania w czasie rzeczywistym (Runtime Monitoring) ### Cyfrowe pochodzenie i walka z dezinformacją Kluczową technologią obronną staje się **Digital Provenance** (Cyfrowe Pochodzenie) oraz standardy C2PA, które pozwalają kryptograficznie poświadczyć autentyczność i źródło treści multimedialnych. Rozwiązania takie jak Google SynthID, Adobe C2PA, Microsoft GUID umożliwiają identyfikację treści generowanych przez AI. ### Zagrożenie kwantowe i kryptografia post-kwantowa Scenariusze "Harvest Now, Decrypt Later" stają się realne - dane szyfrowane dziś mogą być odszyfrowane przez komputery kwantowe w przyszłości. Polska wdraża projekty **kryptografii post-kwantowej (PQC)** oparte na algorytmach Kyber, Dilithium, Falcon i SPHINCS+. ## No-Code/Low-Code: developerzy obywatele przejmują stery W 2026 roku 70-75% nowych aplikacji enterprise będzie zawierało komponenty **no-code lub low-code**, w porównaniu do 25% w 2023 roku. To 3-krotny wzrost w ciągu pięciu lat. ### Dlaczego to się dzieje? - **Szybkość**: skrócenie czasu development o 50-70% - **Koszt**: redukcja kosztów o 40-60% - **Demokratyzacja**: umożliwienie tworzenia rozwiązań przez "citizen developers" - pracowników biznesowych bez znajomości kodowania Kluczowa zmiana w 2026 roku to **AI-assisted development**. Platformy takie jak Microsoft Power Platform pozwalają generować logikę aplikacji, workflows i połączenia danych z promptów w języku naturalnym. ### Wyzwanie: governance na skalę Biggest challenge: utrzymanie jakości, bezpieczeństwa i zgodności bez developer gatekeeping. Leading organizations wdrażają **"governed citizen development"** - frameworki nadzoru umożliwiające szybkie innowacje przy zachowaniu standardów bezpieczeństwa, compliance i spójności architektonicznej. ## Ekonomia AI: weryfikacja ROI i nowe modele cenowe Rok 2026 przyniesie "otrzeźwienie" na rynku. Po latach entuzjazmu, **inwestorzy i zarządy będą domagać się twardych dowodów na zwrot z inwestycji**. ### Koniec ery "AI tourism" Forrester przewiduje, że 25% budżetów na AI zostanie przesuniętych na 2027 rok z powodu opóźnień w weryfikacji wartości. CIO będą zmuszeni ratować projekty AI, które poniosły porażkę z braku kompetencji technicznych. W 2026 liczyć się będą tylko rozwiązania przynoszące **mierzalną poprawę efektywności, redukcję kosztów lub wzrost przychodów**. IDC prognozuje, że 70% CEO z listy G2000 będzie koncentrować ROI z AI na wzroście przychodów, a nie tylko na redukcji zatrudnienia. ### Zmiana modeli cenowych SaaS Tradycyjny model seat-based pricing staje się przestarzały w świecie, w którym pracę wykonują **cyfrowi agenci, a nie ludzie logujący się do systemu**. IDC przewiduje, że do 2028 roku 70% dostawców oprogramowania będzie musiało przebudować modele cenowe, przechodząc na rozliczenia: - **Outcome-based**: płatność za wyniki biznesowe - **Consumption-based**: płatność za zużycie zasobów - **Agent-based**: płatność za liczbę aktywnych agentów AI ## Rynek pracy: redefinicja ról, nie eliminacja AI w 2026 roku nie eliminuje zawodów hurtowo - redefiniuje role zawodowe i wymaga nowych kompetencji. ### Nowe role w erze AI | Nowa rola zawodowa | Kluczowe kompetencje | Zastosowanie | | ------------------ | ---------------------------------------- | ------------------------ | | AI Product Owner | Zarządzanie cyklem życia AI, compliance | Wdrażanie modeli AI | | AI Risk Officer | Zarządzanie ryzykiem, etyka, audyt | Monitorowanie incydentów | | AI Orchestrator | Koordynacja agentów, integracja systemów | Automatyzacja procesów | | Prompt Engineer | Optymalizacja interakcji z modelami | Wszystkie działy | | Data Curator | Zarządzanie danymi treningowymi | AI/ML teams | ### Najbardziej zagrożone vs. odporne stanowiska **Zagrożone**: rutynowe zadania poznawcze - wprowadzanie danych, podstawowe kodowanie, administracja, obsługa klienta pierwszego poziomu. **Odporne**: prace wymagające złożonego osądu, empatii, kreatywności i głębokiej wiedzy dziedzinowej. ### Przekwalifikowanie jako strategia Amazon wdraża program Career Choice, a World Economic Forum promuje inicjatywy Human-Machine Collaboration. Firmy, które sukces odniosą w 2026 roku, traktują transformację jako **intencjonalne reskilling, a nie reaktywną eliminację stanowisk**. ## Polska 2026: gdzie jesteśmy i dokąd zmierzamy ### Strategia cyfrowa i fabryki AI Ministerstwo Cyfryzacji finalizuje plany na lata 2026-2027, kluczowe projekty: - **mObywatel**: aplikacja jako centralny hub usług państwa - **e-Doręczenia**: pełne wdrożenie cyfrowej korespondencji urzędowej - **Fabryki AI**: uruchomienie centrów obliczeniowych w Poznaniu i Krakowie wspierających polskich naukowców i MŚP - **Gigafabryka AI**: projekt klastra ośrodków wiodących (Poznań, Kraków, Wrocław, Warszawa, Gdańsk) - inwestycja 5 mld zł, 2 mld zł z funduszy publicznych ### Luka kompetencyjna i adopcja Mimo postępów, Polska nadrabia zaległości. Wskaźniki adopcji chmury i zaawansowanej analityki w MŚP pozostają poniżej średniej UE. Polski Instytut Ekonomiczny wskazuje, że w 2025 roku zaledwie **8,7% firm stosowało AI** - rok 2026 wymaga gigantycznego wysiłku edukacyjnego. ### Cyberbezpieczeństwo w kontekście geopolitycznym Ze względu na położenie geopolityczne, Polska pozostaje na pierwszej linii cyberwojny. W 2026 eksperci przewidują intensyfikację ataków na infrastrukturę krytyczną oraz kampanii dezinformacyjnych sterowanych przez AI. **Ochrona cyfrowych granic staje się równie ważna jak ochrona granic fizycznych**. ## Praktyczne rekomendacje: co robić w 2026 roku? Po przeanalizowaniu setek stron raportów i prognoz, wyciągam pięć kluczowych rekomendacji dla liderów technologicznych i decision makerów: ### 1. Audyt gotowości agentycznej Nie wdrażaj AI "na siłę". Zamiast tego: - Zidentyfikuj procesy nadające się do autonomizacji (powtarzalne, jasno zdefiniowane, mierzalne) - Przygotuj infrastrukturę danych (Data Governance, jakość danych, dostępność) - Zdefiniuj punkty kontroli i mechanizmy eskalacji do człowieka - Ustal KPI dla wdrożeń AI - bez mierzalnego ROI, projekt nie ma sensu ### 2. Infrastruktura: hybrydowa, a nie monolityczna - **Cloud**: elastyczność dla treningu i eksperymentów - **On-premises**: stabilność dla krytycznego wnioskowania i zgodność z suwerennością danych - **Edge**: ultra-niska latencja dla aplikacji real-time Rozważ wyposażenie pracowników w AI PC - w długim okresie tańsze niż subskrypcje chmurowe liczone od zapytania. ### 3. Compliance jako przewaga konkurencyjna Potraktuj EU AI Act nie jako przeszkodę, ale jako **ramę do budowy bezpiecznego biznesu**. Transparentność przyciągnie klientów zmęczonych dezinformacją i nieprzewidywalnością AI. - Rozpocznij inwentaryzację systemów AI w organizacji - Sklasyfikuj je według poziomu ryzyka (AI Act categories) - Wdróż mechanizmy dokumentacji technicznej i logowania decyzji - Rozważ skorzystanie z piaskownicy regulacyjnej ### 4. Specjalizacja modeli nad uniwersalnością Jeśli działasz w branży regulowanej (finanse, medycyna, prawo), **inwestuj w DSLM zamiast walczyć z uniwersalnymi LLM**. Architektura hybrydowa (foundation model + domain modules) to sweet spot między elastycznością a precyzją. ### 5. Reskilling jako strategia, nie taktyka W Polsce szczególnie istotne: - Zainwestuj w utrzymanie i rozwój pracowników 50+ - ich doświadczenie + nowe narzędzia AI to klucz do stabilności - Stwórz programy upskillingu dla zespołów - AI nie zastąpi ekspertów, ale eksperci bez AI będą zastąpieni przez ekspertów z AI - Zbuduj kulturę eksperymentowania i uczenia się w organizacji ## Podsumowanie: budujemy fundamenty, a nie zabawki Rok 2026 to czas, w którym technologia przestaje być "magią", a staje się "inżynierią". Fascynacja możliwościami AI ustępuje miejsca ciężkiej pracy nad jej wdrożeniem, zabezpieczeniem i skalowaniem. **Kluczowe wnioski**: 1. **Agentic AI** redefiniuje automatyzację - od narzędzi wspomagających do systemów działających autonomicznie 2. **Specjalizacja modeli (DSLM)** wygrywa nad uniwersalnością w branżach regulowanych 3. **Infrastruktura hybrydowa** (cloud + on-premises + edge) to nowy standard 4. **EU AI Act** (sierpień 2026) wymusza transparentność i accountability 5. **ROI i weryfikacja wartości** stają się kluczowe - koniec ery eksperymentów bez strategii 6. **Physical AI** wkracza do produkcji - roboty humanoidalne, autonomiczne fabryki 7. **Cyberbezpieczeństwo prewencyjne** i platformy AI Security to konieczność 8. **Rynek pracy** ewoluuje - nowe role, redefinicja kompetencji, reskilling Wygrają ci, którzy zamiast czekać na opadnięcie kurzu, zaczną budować fundamenty pod nową, autonomiczną rzeczywistość już dziś. A jeśli pytasz się, od czego zacząć - **zacznij od audytu**. Sprawdź, gdzie w Twojej organizacji AI może przynieść mierzalną wartość, jakie procesy nadają się do autonomizacji, jakie dane masz do dyspozycji i czy Twoja infrastruktura jest gotowa. Bo rok 2026 nie będzie o tym, kto ma najlepszą prezentację AI. Będzie o tym, kto ma najlepsze wdrożenie.

Potrzebujesz wsparcia we wdrażaniu AI?

Pomogę Ci ocenić gotowość Twojej organizacji, zidentyfikować procesy do automatyzacji i zaplanować pierwsze kroki.

Umów bezpłatną konsultację
**Źródła i referencje**: Artykuł powstał na podstawie analizy raportów Gartner, Forrester, IDC, Deloitte, McKinsey, Cisco oraz dokumentacji EU AI Act. Wszystkie prognozy i dane są aktualne na dzień 1 stycznia 2026 roku. ## FAQ
### Czym różni się Agentic AI od modeli generatywnych takich jak ChatGPT czy Claude? Agentic AI nie tylko odpowiada na pytania, ale samodzielnie planuje i wykonuje zadania w systemach biznesowych. Zamiast czekać na instrukcje krok po kroku, agenci autonomicznie podejmują decyzje i używają narzędzi, działając jako "wirtualni pracownicy" a nie tylko asystenci.
### Jakie obowiązki nakłada EU AI Act na firmy od sierpnia 2026 roku? Od 2 sierpnia 2026 firmy muszą wdrożyć systemy zarządzania ryzykiem, nadzór ludzki oraz pełną dokumentację techniczną dla systemów AI wysokiego ryzyka. Wymagana jest gwarancja jakości danych treningowych i procedury raportowania incydentów, a kary za brak zgodności sięgają 35 mln euro lub 7% obrotu.
### Dlaczego modele specyficzne dla domeny (DSLM) są lepsze dla biznesu niż ogólne LLM? Modele DSLM oferują wyższą precyzję w specjalistycznych zadaniach (np. prawo, medycyna) przy znacznie niższych kosztach operacyjnych dzięki mniejszej liczbie parametrów. Gwarantują też lepszą zgodność z regulacjami i bezpieczeństwo danych, co jest kluczowe w branżach regulowanych, gdzie ogólne modele często "halucynują".
### Czy AI w 2026 roku zastąpi miejsca pracy czy zmieni ich charakter? AI w 2026 roku nie eliminuje zawodów hurtowo, lecz redefiniuje role, automatyzując rutynowe zadania poznawcze. Powstają nowe stanowiska jak AI Orchestrator czy AI Risk Officer, a kluczem do utrzymania zatrudnienia staje się reskilling i umiejętność współpracy z agentami cyfrowymi.
### Na czym polega przewaga modeli rozumowania (Reasoning Models) nad tradycyjnymi modelami językowymi? Modele rozumowania (System 2) poświęcają dodatkowy czas obliczeniowy na "namysł", symulację i weryfikację hipotez przed udzieleniem odpowiedzi. Pozwala to na rozwiązywanie złożonych problemów logicznych i matematycznych z dokładnością powyżej 90%, kosztem dłuższego czasu reakcji i wyższego zużycia tokenów.
--- # Vibe Coding - jak tworzyć UI z AI bez znajomości designu Source: https://pawel.lipowczan.pl/blog/vibe-coding-przewodnik Published: 2025-12-24 # Vibe Coding - jak tworzyć UI z AI bez znajomości designu Kilka tygodni temu trafiłem na materiał zespołu PageAI o vibe codingu i muszę przyznać - zrobili świetną robotę. Co więcej, ich podejście idealnie pokrywa się z tym, jak sam pracuję z AI przy tworzeniu interfejsów. Jeśli nie masz wykształcenia w UX/UI (jak ja), ale chcesz tworzyć nowoczesne, atrakcyjne interfejsy - ten przewodnik jest dla Ciebie. ## Co to jest Vibe Coding? Vibe coding to podejście do tworzenia interfejsów, gdzie zamiast precyzyjnych mockupów i pixel-perfect designów, **opisujesz wrażenie i vibe**, jaki chcesz osiągnąć. AI (w połączeniu z dobrymi narzędziami) tłumaczy to na działający kod. To demokratyzacja tworzenia UI. AI świetnie radzi sobie z odwzorowywaniem gotowych szablonów, designu ze zdjęcia czy opisów tekstowych. Nawet bez umiejętności projektowych możesz stworzyć coś, co wygląda profesjonalnie. ## 3 Filary Skutecznego Vibe Codingu Z doświadczenia wiem, że vibe coding działa wtedy, gdy masz trzy rzeczy: ### 1. Dobrą bazę (Starter Kit) Nie zaczynaj od zera. Użyj sprawdzonego boilerplate'u, który ma: - Skonfigurowane narzędzia (Vite, React, Tailwind) - Bazowy system designu (kolory, typografia, spacing) - Podstawowe komponenty **Polecam:** [Vibe Coding Starter Kit od PageAI](https://github.com/PageAI-Pro/vibe-coding-starter) To solidna baza z React + Vite + Tailwind + shadcn/ui. Wszystko gotowe do szybkiego startu. ### 2. Dobre i sprawdzone prompty To jest sedno vibe codingu. Musisz wiedzieć **jak rozmawiać z AI**, żeby dostawać sensowne wyniki. Poniżej znajdziesz sprawdzone prompty, które możesz kopiować i dostosowywać. ### 3. Dobry kontekst dla AI AI musi rozumieć: - Jaki jest Twój design system (kolory, czcionki, spacing) - Jakich bibliotek używasz (Tailwind, shadcn/ui, etc.) - Jaki styl projektujesz (minimalistyczny, bold, glassmorphic) Bez kontekstu dostaniesz generyczny, "AI-owaty" design. Z kontekstem - coś unikalnego. ## Krok 1: Zdefiniuj Design Brief Zanim poprosisz AI o kod, musisz wiedzieć **czego chcesz**. Użyj tego promptu: ```text I need help creating a design brief for [type of project]. Please help me define: 1. Design Style & Aesthetic - What visual style best suits this project? - What mood/feeling should the design evoke? - Any specific design movements or styles to reference? 2. Target Audience - Who will use this? - What are their preferences and expectations? - What devices will they primarily use? 3. Key Features & Priorities - What are the core features to highlight? - What's the primary user action? - What should stand out most? Based on my answers, create a comprehensive design brief that I can use with AI coding assistants. ``` **Przykład użycia:** Ja: "I need help creating a design brief for a personal portfolio website." AI pomoże Ci przemyśleć styl (np. minimalistyczny z neonowymi akcentami), grupę docelową (rekruterzy tech, klienci freelance) i kluczowe elementy (hero section, projekty, kontakt). Efekt? Jasny brief, który możesz przekazać w kolejnych promptach. ## Krok 2: Stwórz AI Design Style Reference Tu zaczyna się magia. Stwórz dokument, który AI będzie używał jako punkt odniesienia dla **całego projektu**. ### Visual Style Zdefiniuj styl wizualny w tabeli: | Aspekt | Wybór | Uzasadnienie | |--------|-------|--------------| | **Overall Aesthetic** | Minimalist Brutalism | Podkreśla treść, nowoczesny vibe | | **Color Palette** | Monochromatic + neon accent | Czytelny, wyrazisty, tech-forward | | **Typography** | Inter (sans-serif) | Czytelny, profesjonalny | | **Spacing** | Generous whitespace | Ułatwia fokus, premium feel | | **Visual Weight** | Bold typography, subtle UI | Treść > dekoracje | ### Layout Structures | Typ layoutu | Kiedy używać | Charakterystyka | |-------------|--------------|-----------------| | **Hero Full-Screen** | Landing pages | Maksymalny impact, CTA above the fold | | **Sidebar Layout** | Dashboardy, aplikacje | Nawigacja zawsze widoczna | | **Card Grid** | Portfolio, galerie | Skanowalne, responsywne | | **Split Screen** | Porównania, dual content | 50/50 attention split | ### Color Themes | Theme | Primary | Background | Text | Accent | |-------|---------|------------|------|--------| | **Dark Mode** | `#0f172a` | `#020617` | `#f1f5f9` | `#22d3ee` | | **Light Mode** | `#ffffff` | `#f8fafc` | `#0f172a` | `#0ea5e9` | | **Neon Dark** | `#000000` | `#0a0a0a` | `#ffffff` | `#00ff9d` | **Pro tip:** Zapisz te tabele jako osobny plik (np. `design-reference.md`) i załączaj je w kontekście AI (Cursor, Claude, ChatGPT). ## Krok 3: Implementuj z AI Teraz masz brief i style reference. Czas na kod. Użyj tego promptu: ```text I need you to build [specific component/page] using React and Tailwind CSS. DESIGN CONTEXT: [Wklej swój design brief i style reference] TECHNICAL REQUIREMENTS: - Use React functional components with hooks - Use Tailwind CSS for all styling (no custom CSS) - Make it fully responsive (mobile-first) - Follow accessibility best practices (WCAG 2.1 AA) - Use semantic HTML SPECIFIC REQUIREMENTS FOR THIS COMPONENT: - [Opisz dokładnie co ma robić] - [Jakie sekcje/elementy zawierać] - [Jakieś specjalne interakcje] VISUAL DETAILS: - Follow the [specific style] from the design reference - Use [specific color theme] - Ensure [specific spacing/typography rules] Please provide: 1. Complete, production-ready component code 2. Brief explanation of design decisions 3. Any additional dependencies needed Start with the code, then explain. ``` ### Przykład rzeczywistego użycia: ```text I need you to build a Hero section for a SaaS landing page using React and Tailwind CSS. DESIGN CONTEXT: - Style: Minimalist with gradient background - Color: Dark mode with cyan accent (#22d3ee) - Typography: Inter, bold headlines - Spacing: Generous (p-8, gap-8) TECHNICAL REQUIREMENTS: - React functional component - Tailwind CSS only - Mobile-first responsive - WCAG 2.1 AA compliant SPECIFIC REQUIREMENTS: - Headline + subheadline + CTA button - Gradient background (slate to cyan) - CTA button with hover effect - Centered content, max-width container VISUAL DETAILS: - Use Dark Mode theme from reference - Cyan accent for CTA - Bold typography (font-bold, text-5xl for headline) Start with the code. ``` AI da Ci gotowy komponent, który możesz od razu użyć. ## Design Tokens - Twój System Designu w JSON Jedną z najlepszych praktyk jest trzymanie design tokens w JSON. Dzięki temu AI ma **pojedyncze źródło prawdy** o Twoich kolorach, spacingu, typografii. Przykład `design-tokens.json`: ```json { "colors": { "primary": { "50": "#f0fdfa", "100": "#ccfbf1", "500": "#14b8a6", "900": "#134e4a" }, "neutral": { "50": "#f8fafc", "900": "#0f172a" }, "accent": { "cyan": "#22d3ee", "neon": "#00ff9d" } }, "typography": { "fontFamily": { "sans": ["Inter", "system-ui", "sans-serif"], "mono": ["Fira Code", "monospace"] }, "fontSize": { "xs": "0.75rem", "sm": "0.875rem", "base": "1rem", "lg": "1.125rem", "xl": "1.25rem", "2xl": "1.5rem", "3xl": "1.875rem", "4xl": "2.25rem", "5xl": "3rem" }, "fontWeight": { "normal": "400", "medium": "500", "semibold": "600", "bold": "700" } }, "spacing": { "xs": "0.5rem", "sm": "1rem", "md": "1.5rem", "lg": "2rem", "xl": "3rem", "2xl": "4rem" }, "borderRadius": { "none": "0", "sm": "0.25rem", "md": "0.5rem", "lg": "1rem", "full": "9999px" } } ``` **Jak to użyć?** 1. Stwórz plik `design-tokens.json` w swoim projekcie 2. Zaimportuj go w Tailwind config (`tailwind.config.js`) 3. Załączaj w promptach: "Use colors and spacing from design-tokens.json" AI będzie trzymać się Twojego systemu designu. ## Narzędzia, których używam ### Cursor (mój wybór nr 1) Cursor to IDE oparte na VS Code z wbudowanym AI. Obsługuje: - **Composer** - tworzy/modyfikuje wiele plików jednocześnie - **Chat** - rozmowa z kontekstem całego projektu - **Cmd+K** - inline edycja kodu **Dlaczego Cursor?** Bo AI widzi cały projekt. Wie, jakich komponentów używasz, jaki masz design system, jaką strukturę folderów. Generuje spójny kod. ### Shadcn/ui Biblioteka komponentów, którą możesz instalować przez CLI. Komponenty są **Twoje** - kopiujesz je do projektu i modyfikujesz. AI świetnie zna shadcn/ui, więc jak powiesz "use shadcn Button component", dostaniesz dokładnie to. ### v0.dev (opcjonalnie) Narzędzie Vercel do generowania UI z promptów tekstowych. Dobre na szybkie prototypy, ale kod wymaga często poprawek. ## Praktyczne wskazówki ### 1. Zacznij od małych komponentów Nie próbuj od razu generować całej strony. Zacznij od: - Button - Card - Input - Navigation Gdy masz bazowe komponenty, AI łatwiej komponuje je w większe layouty. ### 2. Iteruj na żywym kodzie Vibe coding to nie "wygeneruj i done". To iteracja: 1. Wygeneruj pierwszą wersję 2. Zobacz w przeglądarce 3. Popraw prompt ("Make the button larger, add hover effect") 4. Wygeneruj ponownie 5. Powtarzaj ### 3. Używaj screenshotów jako input AI (szczególnie Claude z vision) potrafi odwzorować design ze zdjęcia. Znajdź inspirację na Dribbble/Behance, zrób screenshot, wklej do AI: ```text Recreate this design using React and Tailwind CSS. Maintain the layout and color scheme, but adapt it to our design system. ``` Działa zadziwiająco dobrze. ### 4. Trzymaj dokumentację obok Załączaj w kontekście: - `design-reference.md` - Twój style guide - `design-tokens.json` - System designu - `component-patterns.md` - Jak używasz swoich komponentów AI będzie spójne z Twoim projektem. ## Pułapki, których unikaj ### Generyczny "AI design" Bez kontekstu AI generuje nudne, generyczne UI. Zawsze: - Podaj konkretny styl (minimalist, brutalist, glassmorphic) - Wskaż inspiracje (np. "like Stripe's homepage") - Zdefiniuj kolory i typografię ### Za duże komponenty Nie proś AI o "całą landing page". Podziel na: - Hero section - Features section - Pricing section - Footer Mniejsze kawałki = lepsza kontrola. ### Brak design systemu Jeśli każdy komponent ma inne odcienie niebieskiego i inne zaokrąglenia, Twój UI wygląda jak Frankenstein. Stwórz design tokens i trzymaj się ich. ## Kluczowe wnioski 1. **Vibe coding działa** - nawet bez umiejętności designu możesz tworzyć profesjonalne UI 2. **3 filary to klucz** - dobra baza, dobre prompty, dobry kontekst 3. **AI to narzędzie, nie magiczna różdżka** - potrzebujesz jasnego briefu i iteracji 4. **Design system to fundament** - stwórz go raz, używaj wszędzie 5. **Inspiruj się, nie kopiuj** - AI pomaga Ci stworzyć coś unikalnego, nie generycznego ## Zasoby do dalszej nauki - [Vibe Coding Starter Kit](https://github.com/PageAI-Pro/vibe-coding-starter) - solidna baza do startu - [Tailwind CSS Docs](https://tailwindcss.com) - oficjalna dokumentacja - [Shadcn/ui](https://ui.shadcn.com) - biblioteka komponentów - [Realtime Colors](https://realtimecolors.com) - generator palet kolorów - [Font Pair](https://fontpair.co) - inspiracje dla typografii ## Co robić dalej? Jeśli chcesz zacząć z vibe coding: 1. **Sklonuj starter kit** z GitHub 2. **Stwórz design brief** dla swojego projektu (użyj promptu z tego artykułu) 3. **Zdefiniuj style reference** - tabele ze stylem wizualnym 4. **Wygeneruj pierwszy komponent** (zacznij od Button) 5. **Iteruj i ucz się**

Potrzebujesz wsparcia we wdrożeniu AI w rozwoju produktu?

Pomogę Ci skonfigurować AI tools, zautomatyzować workflow i wdrożyć vibe coding w Twoim projekcie. Od wyboru narzędzi przez konfigurację środowiska po optymalizację procesów deweloperskich.

Umów bezpłatną konsultację
## FAQ
### Co to jest vibe coding i dla kogo jest przeznaczony? Vibe coding to podejście do tworzenia UI, gdzie zamiast pixel-perfect mockupów opisujesz wrażenie i styl, jaki chcesz osiągnąć. AI tłumaczy to na działający kod. Przeznaczony dla osób bez wykształcenia UX/UI, które chcą tworzyć profesjonalne interfejsy - developerów, przedsiębiorców, twórców produktów.
### Jakie są 3 filary skutecznego vibe codingu? Dobra baza (starter kit z Vite, React, Tailwind i podstawowymi komponentami), dobre prompty (konkretne instrukcje dla AI z kontekstem designu) i dobry kontekst (design system, biblioteki, styl projektu). Bez tych trzech elementów AI generuje generyczny, niespójny design.
### Dlaczego warto trzymać design tokens w pliku JSON? Design tokens to pojedyncze źródło prawdy o kolorach, spacingu i typografii projektu. AI załączając ten plik w kontekście trzyma się spójnego systemu designu. Bez tokenów każdy komponent ma inne odcienie i zaokrąglenia - interfejs wygląda jak Frankenstein złożony z przypadkowych elementów.
### Jakie narzędzia są najlepsze do vibe codingu z AI? Cursor (IDE z wbudowanym AI widzi cały projekt), shadcn/ui (komponenty kopiowane do projektu, świetnie znane przez AI) i opcjonalnie v0.dev (Vercel do szybkich prototypów). Cursor to wybór numer jeden - AI generuje spójny kod, bo zna strukturę folderów i istniejące komponenty.
### Jak uniknąć generycznego "AI design" przy vibe codingu? Zawsze podaj konkretny styl (minimalist, brutalist, glassmorphic), wskaż inspiracje ("jak strona Stripe"), zdefiniuj kolory i typografię w design reference. Dziel zadania na małe komponenty zamiast prosić o całą stronę naraz. Iteruj - vibe coding to proces generowania, podglądu i poprawiania, nie jednorazowe wygenerowanie gotowego UI.
--- # Dane jako paliwo biznesu - od Excela do AI Source: https://pawel.lipowczan.pl/blog/dane-jako-paliwo-biznesu Published: 2025-12-21 Dane to nowa ropa naftowa. To stwierdzenie pojawia się w branży technologicznej od 2006 roku, kiedy matematyk Clive Humby po raz pierwszy spopularyzował to porównanie. I tak jak ropa naftowa, dane same w sobie nie mają wartości - dopiero po odpowiedniej "rafinacji" stają się paliwem napędzającym biznes. Z mojego 15-letniego doświadczenia jako programista i 4 lat pracy z no-code wiem jedno: **większość firm siedzi na kopalni złota, ale nie potrafi go wydobyć**. Mają Tony danych, ale brakuje im narzędzi i procesów, żeby przekształcić je w wartościowe wnioski. A teraz, w 2025 roku, stoimy na progu kolejnej rewolucji. Agenty AI takie jak Claude Code, GitHub Copilot czy Cursor zmieniają zasady gry - to, co wcześniej wymagało wielomiesięcznych projektów IT, dziś można zbudować w dni. Połączenie zwinności no-code z mocą tradycyjnego kodu wspieranego przez AI otwiera zupełnie nowe możliwości. Dzisiaj pokażę wam, jak przejść tę drogę - od chaosu w Excelach do uporządkowanego systemu danych, który faktycznie pracuje dla biznesu. ## Problem: Topimy się w danych, umieramy z pragnienia informacji ### 180 zettabajtów chaosu Wyobraźcie sobie **180 zettabajtów** danych. To ilość informacji, która zostanie wygenerowana w 2025 roku. Żeby to sobie uzmysłowić: gdybyśmy zapisali te dane na kartkach A4 i ułożyli jedną na drugiej, moglibyśmy polecieć do Księżyca i z powrotem **60 razy**. Albo potrzebowalibyśmy **62 miliardy pendrive'ów po 16GB**. Ilość danych podwaja się co 3 lata. I prawdopodobnie będzie się podwajać jeszcze szybciej przez rozwój AI i IoT. **Paradoks jest taki**: mamy mnóstwo danych, ale nie jesteśmy w stanie wyciągnąć z nich informacji. To jak siedzieć przed spiżarnią pełną składników, ale bez kucharza i książki kucharskiej - nie jesteśmy w stanie przygotować pizzy. ### Cztery główne wyzwania Z doświadczenia pracy z dziesiątkami firm widzę, że problemy powtarzają się jak mantra: **1. Silosy danych** Dane klientów w jednym systemie, zamówienia w drugim, faktury w trzecim. Żadna integracja. Chcesz zobaczyć pełny obraz klienta? Musisz sklecić 3-4 różne raporty i ręcznie je połączyć w Excelu. **2. Niska jakość i niespójność** - Duplikaty (ten sam klient zapisany 5 razy) - Literówki (Kowalski, Kowalsky, Kowalskii) - Różne formaty (2025-12-21 vs 21.12.2025 vs 21/12/25) - Błędne dane (ktoś wpisał datę urodzenia zamiast daty zamówienia) Każda analiza oparta na takich danych prowadzi na manowce. **3. Trudny dostęp** Dane leżą w bazach SQL, do których zwykły pracownik nie ma dostępu. Albo ma dostęp, ale nie potrafi napisać zapytania. Każdy prosty raport wymaga zgłoszenia do IT i tygodnia oczekiwania. **4. Brak kultury danych** Nawet jak dajemy ludziom narzędzia, to nie wiedzą: - Jakie pytania zadawać - Jak interpretować wyniki - Jak wyciągać wnioski - Jak te wnioski wdrażać Badania z 2020 roku pokazują, że mimo dostępu do danych, firmy: - Nie mają wspólnego obrazu sytuacji - Odkładają podejmowanie decyzji (bo "nie mają pewności") - Tracą cenny czas na szukanie informacji - Marnują wartość ekonomiczną i finansową swoich danych ## Droga od surowca do wartości Zanim dane wniosą coś wartościowego, muszą przejść przez kilka kluczowych etapów: 1. **Pozyskiwanie** - zbieranie danych z różnych źródeł 2. **Integracja** - łączenie różnych źródeł w jeden spójny obraz (to najważniejszy krok!) 3. **Czyszczenie i transformacja** - bez tego analizy opierają się na fałszywych przesłankach 4. **Składowanie** - w repozytorium analitycznym, dostępnym dla właściwych osób 5. **Analiza i wizualizacja** - dashboardy, raporty, wykresy 6. **Wdrożenie wniosków** - faktyczne wykorzystanie informacji w podejmowaniu decyzji **Tu dopiero pojawia się wartość biznesowa.** Większość firm utyka już na kroku 2-3. Mają dane, ale nie potrafią ich połączyć i wyczyścić na tyle, żeby mogły być podstawą decyzji. ## Rozwiązanie część 1: No-code i demokratyzacja danych Przez ostatnie 4 lata obserwuję, jak no-code zmienia sposób pracy z danymi w firmach. Platformy takie jak **Airtable, Make, n8n czy Zapier** radykalnie obniżają barierę wejścia. Ludzie z działów biznesowych, marketingu, finansów mogą sami: - Tworzyć struktury danych - Integrować różne źródła - Tworzyć proste raporty i dashboardy - Automatyzować przepływ informacji Nie muszą czekać tygodniami na dział IT. Nie muszą znać SQL, Python czy JavaScript. ### Korzyści demokratyzacji danych **Szybkość**: Prototyp systemu w Airtable można zbudować w godziny, nie miesiące. **Koszt**: Nie trzeba angażować drogich deweloperów do prostych raportów. **Bliskość biznesu**: Ludzie, którzy najlepiej znają dane (bo z nich korzystają), mogą sami budować rozwiązania. **Zaangażowanie**: Pracownicy czują, że mają wpływ - mogą sami rozwiązywać swoje problemy. ### Ale demokratyzacja wymaga równowagi Demokratyzacja danych to nie tylko technologia. To dwie sprzeczne wartości, które muszą współistnieć: **Data literacy** - umiejętności użytkowników: - Efektywne wykorzystanie danych - Zadawanie odpowiednich pytań - Wyciąganie sensownych wniosków - Rozumienie ograniczeń i błędów **Data governance** - zarządzanie danymi: - Nie wszyscy mogą mieć dostęp do wszystkiego - Muszą być procedury i ramy bezpieczeństwa - Trzeba dbać o spójność i jakość - Ktoś musi pilnować, żeby dane nie zostały "popsute" Trzeba szukać **złotego środka** - zarządzanie nie może być hamulcem, ale musi być na tyle zwinne, żeby wspierać użytkowników, a nie ich blokować. ## Rozwiązanie część 2: AI jako game-changer I tu dochodzimy do momentu, w którym jesteśmy teraz - końca 2025 roku. **AI zmienia absolutnie wszystko w kontekście pracy z danymi.** ### Co AI daje już dziś **1. Rozmowa z danymi w języku naturalnym** Nie musisz znać SQL. Piszesz: "Pokaż mi top 10 klientów według wartości zamówień w tym kwartale" - i dostajesz odpowiedź. To radykalnie obniża barierę data literacy. **2. Czyszczenie i normalizacja danych** AI potrafi znaleźć duplikaty, nawet jak są literówki. Potrafi ujednolicić formaty. Potrafi wypełnić braki na podstawie kontekstu. **3. Znajdowanie wzorców** W dużych zbiorach danych AI znajdzie korelacje, których człowiek by nie zauważył. Np. "klienci, którzy kupują produkt X w piątek, częściej wracają po produkt Y w poniedziałek". **4. Generowanie analiz i prognoz** Na podstawie danych historycznych AI może przewidywać przyszłe trendy, popyt, ryzyko churn klientów. **5. Automatyzacja raportowania** Zamiast ręcznie tworzyć raporty co tydzień, AI może generować podsumowania, wykresy i wnioski automatycznie. ### Ale też ryzyka Musimy być świadomi ograniczeń: - **Halucynacje** - AI może wymyślać fakty, które brzmią przekonująco - **Czarna skrzynka** - nie zawsze wiemy, skąd AI wzięło informację - **Niedeterministyczność** - za każdym razem może dać inną odpowiedź - **Garbage in, garbage out** - śmieciowe dane dadzą śmieciowe wnioski - **Prywatność** - dane mogą wyciekać do modeli trenowanych przez zewnętrzne firmy - **Bias** - nawet uporządkowane dane mogą prowadzić do błędnych wniosków przez uprzedzenia w modelu Dlatego AI to narzędzie, nie zamiennik myślenia. ## Przełom: Agenty AI dla programistów A teraz najważniejsze: **nowa generacja narzędzi AI dla deweloperów zmienia całą grę**. Przez lata mieliśmy wybór: - **No-code**: szybko, tanio, ale ograniczone możliwości - **Tradycyjny kod**: nieograniczone możliwości, ale wolno i drogo ### Agenty AI łączą te dwa światy Narzędzia takie jak: - **Claude Code** (którego używam do pisania tego artykułu) - **GitHub Copilot** - **Cursor** - **Windsurf** ...dają programistom **zwinność no-code + moc tradycyjnego kodu**. Co wcześniej było możliwe tylko przez no-code (szybkie prototypowanie, iteracyjne budowanie rozwiązań), teraz programiści z pomocą AI mogą robić równie szybko - ale z pełną elastycznością kodu. ### Przykład z życia Niedawno budowałem system do analizy danych z kilku źródeł (Airtable, API zewnętrzne, pliki CSV). Wcześniej: - W no-code: szybko, ale nie dam rady zrobić złożonych transformacji - W tradycyjnym kodzie: 2-3 tygodnie pracy Z Claude Code: **3 dni**. AI pomogło mi: - Szybko napisać skrypty do integracji API - Wygenerować kod do czyszczenia danych - Stworzyć ETL pipeline - Napisać testy - Zoptymalizować zapytania Nie musiałem pamiętać składni każdej biblioteki. Nie musiałem szukać rozwiązań na Stack Overflow. Po prostu opisywałem, co chcę osiągnąć - a AI generowało kod, który mogłem przejrzeć, zrozumieć i dostosować. ### Co to oznacza dla zarządzania danymi **1. Szybsze wdrażanie rozwiązań** Projekty, które wcześniej trwały miesiące, teraz realizujemy w tygodnie. To oznacza szybszy ROI i możliwość eksperymentowania. **2. Mniejsze bariery wejścia - ale z ostrzeżeniem** AI obniża próg wejścia, ale **nie eliminuje potrzeby doświadczenia**. Junior developer z pomocą AI może zrobić więcej, ale agent AI sam w sobie jest jak junior - może popełniać błędy, nie rozumieć kontekstu biznesowego, generować kod z lukami bezpieczeństwa. **Doświadczenie seniora nie jest wymogiem, ale jest zdecydowanie zalecane.** Ktoś doświadczony musi mieć kontrolę nad tym, co AI generuje - weryfikować podejścia, łapać błędy, oceniać jakość. AI to potężne narzędzie, ale wymaga nadzoru. **3. Lepsza jakość kodu** AI pomaga w pisaniu testów, dokumentacji, optymalizacji. Kod jest czystszy i łatwiejszy w utrzymaniu. **4. Wybór code vs no-code - zależy od zespołu** Wybór podejścia zależy od zasobów w organizacji: **Masz senior programistę?** → Idź w kod z AI. Dostaniesz elastyczność, brak ograniczeń i pełną kontrolę. **Nie masz technicznego zaplecza?** → Zostań przy no-code (Airtable, Make, n8n). To rozwiązanie zdecydowanie bardziej przystępne dla ludzi bez doświadczenia programistycznego. **Hybryda no-code + code** ma sens tylko wtedy, gdy masz na pokładzie kogoś mocnego technicznie. Wtedy to najlepsza opcja: zachowujesz elastyczność kodu, zwinność prototypowania i praktyczny brak ograniczeń, które czasami pojawiają się w narzędziach no-code. **5. Rozmowa z danymi w języku naturalnym - ale z mocą SQL** Nie musisz wybierać między prostotą a elastycznością. AI może przetłumaczyć twoje pytanie na złożone zapytanie SQL lub skrypt Pythonowy. ## Case study: Od chaosu do klarowności Teoria to jedno, praktyka to drugie. Pokażę wam, jak przeszliśmy tę transformację w naszej własnej firmie. ### Kontekst 22Ventures był holdingiem składającym się z kilku firm: Automation House, Tigers, Huciao, Sowicki Legal. Automation House powstała właśnie po to, żeby wewnętrznie poukładać procesy i dane - a potem szeroko wprowadzać te rozwiązania na rynek. Zaczęliśmy tam, gdzie większość firm: **chaos w Excelach**. ### Problem 1: Niejednolitość danych Każdy zespół miał swoje Excele. Te same dane wpisywane były różnie: - Duplikaty klientów (ten sam klient 5 razy w bazie) - Literówki (Kowalski, Kowalsky, Kowalskii) - Różne formaty dat, kwot, nazw - Błędy przepisywania **Rozwiązanie**: Normalizacja w Airtable - Stworzenie słowników (lista unikalnych wartości) - Usunięcie duplikatów - Ujednolicenie formatów - Walidacja na poziomie pól (tylko określone wartości, formaty) ### Problem 2: Wiele źródeł danych Każdy dział miał swój Excel: - HR - lista pracowników - Finance - faktury i płatności - Projekty - zadania i timesheets - Sales - leady i klienci Dane były zduplikowane między działami. Brak procedur synchronizacji. Brak kultury organizacji. **Rozwiązanie**: Migracja do Airtable Płynne przeniesienie z Exceli do Airtable. Konsolidacja w jednym workspace z odpowiednimi bazami i relacjami. **Przykład transformacji**: Team leader **W Excelu**: Przy każdym projekcie była pełna linia z danymi team leadera: imię, nazwisko, email, telefon, stawka. Zmiana danych team leadera = zmiana w 20 miejscach ręcznie. **W Airtable**: Baza pracowników + baza projektów połączone relacją. Team leader to linked record. Zmiana danych w jednym miejscu - działa wszędzie automatycznie. ### Problem 3: Brak słowników Dane wpisywane były "z palca". Każdy wpisywał po swojemu. **Przykład**: Benefity pracownicze W Excelu: Przy każdym pracowniku była kolumna "Benefity" z pełnym opisem i ceną. Zmiana ceny karty sportowej = zmiana ręczna przy 50 pracownikach. **W Airtable**: Baza benefitów (słownik) z cenami. Pracownicy mają linked records do benefitów. Zmiana ceny w jednym miejscu - aktualizuje się automatycznie dla wszystkich. ### Problem 4: Nadmiarowość i powielanie Excel z natury prowadzi do powielania danych. Każda zmiana wymaga propagacji ręcznej. **Rozwiązanie**: Single source of truth W Airtable każdy rekord istnieje w jednym miejscu. Relacje łączą dane bez duplikowania. Zmiana w jednym miejscu = zmiana wszędzie. ### Rezultaty transformacji **Przed**: - 15 różnych Exceli - 3-4 godziny tygodniowo na przygotowanie raportów - Błędy w danych przy każdej analizie - Frustracja zespołu - Decyzje oparte na "przeczuciu" **Po**: - Jeden system w Airtable - Raporty generowane automatycznie - Dane czyste i spójne - Zespół zadowolony (ma narzędzia, które działają) - Decyzje oparte na faktach **I co najważniejsze**: Ten system ewoluuje. Zaczęliśmy od poziomu 1-2, teraz jesteśmy na 3-4 z moich 5 poziomów dojrzałości danych. ## 5 poziomów dojrzałości danych Na podstawie pracy z dziesiątkami firm wypracowałem model oceny, na jakim etapie jest organizacja w zarządzaniu danymi: ### Poziom 1: Ad hoc - Excele rozproszone po całej firmie - Pierwsze próby analiz (proste formuły SUM, AVERAGE) - Każdy robi po swojemu - Brak standardów ### Poziom 2: Konsolidacja - Połączenie źródeł danych w jedno miejsce - Np. migracja z Exceli do Airtable - Wszystko w jednym systemie - Podstawowe relacje między danymi ### Poziom 3: Standaryzacja - Procedury gromadzenia danych - Standardy przetwarzania - Data governance (kto ma dostęp do czego) - Słowniki i walidacja - Procesy czyszczenia danych ### Poziom 4: Optymalizacja - Wykorzystanie finansowe danych - Tworzenie produktów opartych o dane - Monetyzacja (sprzedaż danych, insights) - Zaawansowane analizy i prognozy ### Poziom 5: Innowacja - Kultura całkowicie zorientowana na dane - Każda decyzja oparta o fakty - Ustandaryzowane procesy na całej firmie - Regularna monetyzacja - Ciągła poprawa jakości i przepływu danych - Eksperymenty i testowanie hipotez **Większość firm jest na poziomie 1-2. Mało kto dochodzi do 4-5.** Ale z pomocą no-code i AI ten proces można przyspieszyć wielokrotnie. ## Co możesz zrobić dzisiaj ### Krok 1: Oceń swoją dojrzałość Zadaj sobie pytania: - Na ilu systemach/Excelach trzymamy dane? - Jak długo trwa przygotowanie podstawowego raportu? - Jak często znajdujemy błędy w danych? - Czy mamy procedury wprowadzania danych? - Czy ludzie ufają naszym danym? To pozwoli ci określić, na którym poziomie jesteś. ### Krok 2: Zacznij od jednego obszaru Nie próbuj naprawiać wszystkiego naraz. Wybierz jeden problem: - Może to lista klientów z duplikatami - Może to chaos w projektach - Może to proces raportowania sprzedaży Zacznij małym projektem pilotażowym. ### Krok 3: Wybierz odpowiednie narzędzie **Dla prostych przypadków**: Airtable, Notion, Google Sheets z Apps Script **Dla automatyzacji**: Make, n8n, Zapier **Dla zaawansowanych analiz**: Python + Pandas (z pomocą AI), Power BI, Tableau **Dla integracji no-code + code**: Airtable + custom skrypty napisane z pomocą Claude Code lub Cursor ### Krok 4: Zbuduj prototyp z AI Nie musisz być ekspertem. Użyj Claude, ChatGPT lub innego LLM: - Opisz swój problem - Poproś o propozycję struktury danych - Poproś o kod do integracji/czyszczenia - Iteruj i dopracowuj Z pomocą AI możesz zbudować działający prototyp w godziny, nie tygodnie. ### Krok 5: Testuj, ucz się, iteruj Nie szukaj perfekcji od razu. Wdróż coś prostego, zobacz jak działa, zbierz feedback, popraw. **Zwinność to klucz.** ## Kluczowe wnioski ### 1. Dane bez rafinacji to tylko śmieci Nie wystarczy mieć dużo danych. Trzeba je połączyć, wyczyścić, ujednolicić - dopiero wtedy mają wartość. ### 2. No-code demokratyzuje dostęp Airtable, Make, n8n - te narzędzia dają ludziom z biznesu moc budowania rozwiązań bez czekania na IT. ### 3. AI zmienia zasady gry Rozmowa z danymi w języku naturalnym, automatyczne czyszczenie, znajdowanie wzorców - AI obniża barierę wejścia do zaawansowanych analiz. ### 4. Agenty AI łączą no-code i tradycyjny kod Claude Code, GitHub Copilot, Cursor - teraz programiści mogą budować równie zwinnie jak no-code, ale z pełną mocą kodu. To najlepsze z dwóch światów. ### 5. Demokratyzacja wymaga równowagi Data literacy (umiejętności) + data governance (zarządzanie) muszą iść w parze. Łatwy dostęp bez kontroli to chaos. Kontrola bez dostępu to hamulec. ### 6. Zacznij małym krokiem Nie czekaj na wielki projekt transformacji. Wybierz jeden problem, zbuduj prototyp, testuj, ucz się. Małe zwycięstwa budują momentum. ### 7. Jakość > Ilość Lepiej mieć 10 dobrze uporządkowanych, czystych źródeł danych niż 100 chaotycznych Exceli. ## Następne kroki Świat danych zmienia się błyskawicznie. To, co rok temu było science fiction (AI generujące kod, rozmowa z danymi po polsku), dziś jest codziennością. **Firmy, które opanują sztukę przekształcania danych w wartość biznesową, będą liderami w swoich branżach.** A technologia już nie jest barierą. No-code obniżył próg wejścia. AI dał moc analizy. Agenty AI dały programistom zwinność. **Teraz jedyną barierą jest decyzja: zacząć, czy czekać.** Z mojego doświadczenia wiem, że te firmy, które zaczęły rok temu - dziś mają przewagę konkurencyjną. Te, które zaczną dziś - będą ją mieć za rok. A te, które będą czekać... będą tonąć w morzu danych, umierając z pragnienia informacji.

Potrzebujesz pomocy z uporządkowaniem danych w firmie?

Pomogę Ci przejść od chaosu w Excelach do spójnego systemu zarządzania danymi. Od audytu i konsolidacji źródeł przez migrację do Airtable/baz danych po automatyzację raportowania.

Umów bezpłatną konsultację
## FAQ
### Co oznacza stwierdzenie "dane to nowa ropa naftowa" w praktyce biznesowej? Dane, podobnie jak ropa, same w sobie nie mają wartości - muszą przejść "rafinację" (czyszczenie, integrację, analizę), żeby stać się paliwem dla decyzji biznesowych. Większość firm siedzi na kopalni złota, ale nie potrafi go wydobyć - mają tony danych, ale brakuje im procesów i narzędzi do przekształcenia ich w użyteczne informacje.
### Jakie są główne przeszkody w wykorzystaniu danych przez firmy? Cztery kluczowe wyzwania: silosy danych (informacje rozproszone w różnych systemach bez integracji), niska jakość (duplikaty, literówki, niespójne formaty), trudny dostęp (potrzeba IT do każdego raportu) oraz brak kultury danych (ludzie nie wiedzą, jakie pytania zadawać i jak interpretować wyniki).
### Kiedy wybrać rozwiązania no-code, a kiedy kod z pomocą AI do zarządzania danymi? Masz senior programistę w zespole - wybierz kod z AI (Claude Code, Cursor) dla pełnej elastyczności. Nie masz technicznego zaplecza - zostań przy no-code (Airtable, Make, n8n). Hybryda ma sens tylko z mocnym zapleczem technicznym - wtedy łączysz zwinność prototypowania z praktycznym brakiem ograniczeń.
### Jak ocenić poziom dojrzałości danych w firmie? Pięć poziomów: (1) Ad hoc - rozproszone Excele bez standardów, (2) Konsolidacja - dane w jednym systemie, (3) Standaryzacja - procedury, governance, walidacja, (4) Optymalizacja - monetyzacja i zaawansowane analizy, (5) Innowacja - kultura w pełni oparta na danych. Większość firm utyka na poziomie 1-2.
### Od czego zacząć transformację danych w firmie? Wybierz jeden konkretny problem (np. duplikaty klientów, chaos w projektach), zbuduj mały prototyp w Airtable lub z pomocą AI, przetestuj z zespołem, zbierz feedback i iteruj. Nie szukaj perfekcji od razu - małe zwycięstwa budują momentum. Transformacja całej firmy naraz to przepis na porażkę.
--- # Airtable vs Excel - Kiedy warto zmienić arkusz na bazę danych? Source: https://pawel.lipowczan.pl/blog/airtable-vs-excel-migracja Published: 2025-12-19 # Airtable vs Excel - Kiedy warto zmienić arkusz na bazę danych? Przez lata Excel był synonimem organizacji danych w firmach. Każdy go zna, każdy używa. Ale z własnego doświadczenia wiem, że w pewnym momencie klasyczne arkusze kalkulacyjne przestają wystarczać. Gdy zespół rośnie, dane się komplikują, a procesy wymagają lepszej współpracy - wtedy zaczyna się frustracja. Właśnie dlatego coraz więcej firm decyduje się **migrate excel to airtable**. I nie chodzi tu o żaden technologiczny snobizm - to po prostu praktyczne rozwiązanie realnych problemów, które Microsoft Excel ma w DNA. ## Problem: Ograniczenia klasycznych arkuszy Zanim przejdziemy do rozwiązań, spójrzmy szczerze na to, z czym borykają się zespoły używające Excela: **1. Piekło wersjonowania** Znacie to uczucie? `Projekty_v2.xlsx`, `Projekty_v2_final.xlsx`, `Projekty_v2_final_NAPRAWDE_OSTATECZNY.xlsx`. Emaile latają tam i z powrotem, nikt nie wie która wersja jest aktualna, a każdy pracuje na swoim pliku. Efekt? Chaos, duplikacja pracy i błędy. **2. Relacje między danymi to koszmar** Załóżmy, że prowadzisz kalendarz contentu. Masz artykuły, autorów, kampanie marketingowe, statusy publikacji. W Excelu? Mnóstwo VLOOKUP-ów, połączone arkusze, które łamią się przy najmniejszej zmianie struktury. Ręczne kopiowanie danych. Ryzyko błędów na każdym kroku. **3. Współpraca w czasie rzeczywistym? Zapomnij** Tak, wiem - jest Excel Online i Google Sheets. Ale szczerze mówiąc, to wciąż nie jest prawdziwa współpraca. Brak kontroli uprawnień, brak historii zmian, brak elastycznych widoków dla różnych ról w zespole. **4. Wizualizacja i raportowanie wymaga gimnastyki** Chcesz zobaczyć projekty na tablicy Kanban? Kalendarz deadlinów? Galerię z miniaturkami? W Excelu to albo makro, albo osobny dashboard, albo... po prostu nie da się tego zrobić wygodnie. Sam przez lata tworzyłem skomplikowane arkusze dla klientów i za każdym razem docierałem do punktu, gdzie pomyślałem: "Musi być lepszy sposób". I właśnie wtedy poznałem **airtable database builder**. ## Rozwiązanie: Airtable jako relacyjna baza z interfejsem arkusza ### Czym właściwie jest Airtable? Najprościej mówiąc: **Airtable to relacyjna baza danych, która wygląda i działa jak arkusz kalkulacyjny**. To kluczowa różnica w porównaniu **airtable vs excel**. Pod spodem to prawdziwa baza danych z relacjami, typami pól i integritością danych. Ale na wierzchu? Intuicyjny interfejs, który nie wymaga znajomości SQL czy programowania. ### Relacje między tabelami - game changer W Airtable możesz stworzyć: - Tabelę "Artykuły na blogu" - Tabelę "Autorzy" - Tabelę "Kampanie marketingowe" A potem **połączyć je ze sobą**. Kliknięcie, przeciągnięcie i już widzisz wszystkie artykuły danego autora. Wszystkie materiały powiązane z kampanią. Bez VLOOKUP-ów, bez wzorów, bez ryzyka rozsypania się struktury. To nie teoria - używam tego na co dzień w automation.house. Nasza baza wiedzy o klientach, projektach i procesach żyje w Airtable i oszczędza nam kilkadziesiąt godzin miesięcznie. ### Różne widoki dla różnych ról To jedna z moich ulubionych funkcji. Te same dane możesz zobaczyć jako: - **Grid** - klasyczny widok arkusza - **Calendar** - idealne do deadlinów i planowania - **Kanban** - dla zarządzania projektami w stylu Trello - **Gallery** - świetne dla portfolios, produktów, grafik - **Form** - do zbierania danych od zewnętrznych osób Marketing patrzy na kampanie przez Kanban. Content writer przez kalendarz publikacji. Manager przez tabelę z filtrami. **Każdy widzi to, czego potrzebuje.** ## Dlaczego firmy migrują z Excel do Airtable? Z rozmów z klientami i własnego doświadczenia wyłoniło się pięć głównych powodów: ### 1. Prawdziwa współpraca zespołowa Wszyscy pracują na tej samej bazie ("base" w nomenklaturze Airtable) w czasie rzeczywistym. Zmiany są natychmiastowe. Możesz komentować rekordy, oznaczać ludzi, ustawiać przypomnienia. To nie jest już narzędzie do pracy indywidualnej - to platforma zespołowa. ### 2. Automatyzacja powtarzalnych zadań Airtable ma wbudowany system automatyzacji. Przykłady z życia wziętego: - Gdy status projektu zmienia się na "Do akceptacji" → wyślij powiadomienie do managera - Gdy deadline mija za 3 dni → wyślij email do odpowiedzialnej osoby - Gdy nowy rekord zostaje dodany przez formularz → stwórz zadania w połączonych tabelach W Excelu? Potrzebujesz VBA, makr albo zewnętrznych narzędzi. W Airtable? Klikasz i konfigurujesz. ### 3. Integracje z resztą stacku Airtable świetnie łączy się z: - Slack (powiadomienia) - Gmail (wysyłanie emaili) - Zapier/Make (zaawansowane workflow) - Google Calendar (synchronizacja eventów) - I setkami innych narzędzi To sprawia, że Airtable staje się centralnym hubem danych dla całej firmy. ### 4. Kontrola uprawnień i bezpieczeństwo Możesz precyzyjnie określić, kto ma dostęp do jakiej tabeli, kto może edytować, a kto tylko czytać. Możesz ukrywać pola przed określonymi rolami. Historia zmian pokazuje, kto i kiedy coś modyfikował. ### 5. Skalowalność bez bólu Zacznij od prostej bazy z 3 tabelami. Potem dodaj kolejne. Połącz je relacjami. Dodaj automatyzacje. Stwórz publiczne formularze. Zbuduj interfejsy dla klientów. Airtable rośnie razem z Twoimi potrzebami, a nie przeciwko nim - jak to często bywa z rozrośniętymi arkuszami Excela. ## Jak przenieść dane z Excel do Airtable? Dobrze, przekonałem Cię. Ale jak faktycznie **migrate excel to airtable**? ### Krok 1: Przygotuj dane w Excelu - Upewnij się, że każda tabela ma nagłówki w pierwszym wierszu - Usuń puste wiersze i kolumny - Rozdziel dane logicznie (jeśli masz kilka "podmiotów" w jednym arkuszu, rozważ podział) ### Krok 2: Import do Airtable Airtable pozwala na import plików `.xlsx` i `.csv`. To proste: 1. Stwórz nową bazę w Airtable 2. Kliknij "Add or import" → "CSV or TSV" 3. Wgraj plik 4. Airtable automatycznie rozpozna typy kolumn (tekst, liczby, daty) ### Krok 3: Dopracuj strukturę Tutaj dzieje się magia. Po imporcie: - Zmień typy pól tam, gdzie Airtable się pomylił (np. zmień text na email, URL, telefon) - Rozdziel dane na osobne tabele (np. oddziel klientów od projektów) - Stwórz **relacje między tabelami** używając pola "Link to another record" - Usuń duplikaty i uporządkuj dane ### Krok 4: Stwórz widoki - Zbuduj widoki Calendar dla dat - Kanban dla statusów - Gallery dla projektów z obrazkami - Odfiltrowane widoki dla konkretnych zespołów ### Krok 5: Dodaj automatyzacje Zacznij od prostych: - Powiadomienie Slack gdy nowy rekord - Email gdy deadline się zbliża - Automatyczna zmiana statusu ## Airtable vs Excel - dla kogo Airtable? Nie twierdzę, że Excel jest zły. Ma swoje miejsce. Ale Airtable jest lepszy dla: **✅ Zespołów współpracujących nad danymi** Excel jest dla indywidualnej pracy, Airtable dla zespołowej. **✅ Danych z relacjami i zależnościami** Klienci → Projekty → Faktury → Płatności. W Airtable to naturalne, w Excelu - ból. **✅ Procesów wymagających różnych widoków** Kalendarz, Kanban, Grid, Gallery - wszystko z tych samych danych. **✅ Automatyzacji i workflow** Airtable ma to wbudowane, Excel wymaga VBA lub zewnętrznych narzędzi. **✅ Integracji z innymi narzędziami** API, Zapier, Make - Airtable jest stworzony do łączenia się z ekosystemem. Z kolei **Excel wciąż wygrywa** przy: - Zaawansowanych obliczeniach finansowych i statystycznych - Indywidualnej analizie danych - Jednorazowych raportach - Pracy offline bez dostępu do internetu ## Co możesz zrobić dzisiaj? Jeśli zastanawiasz się nad migracją, zacznij małymi krokami: 1. **Wybierz jeden proces/arkusz** do przetestowania w Airtable 2. **Skorzystaj z darmowego planu** Airtable (wystarczy na start) 3. **Zaimportuj dane** i pobaw się strukturą przez tydzień 4. **Sprawdź czy rozwiązuje Twoje problemy** z Excelem 5. **Dopiero wtedy rozważaj pełną migrację** Z mojego doświadczenia najlepiej sprawdzają się jako pierwsze: - Kalendarze contentu - Bazy klientów/CRM - Zarządzanie projektami - Listy zadań zespołowych - Inwentarz/katalogi produktów ## Kluczowe wnioski Porównanie **airtable vs excel** nie ma jednoznacznego zwycięzcy - to zależy od kontekstu. Ale jeśli: - Pracujesz w zespole - Dane mają złożone relacje - Potrzebujesz różnych widoków tych samych danych - Chcesz automatyzować workflow - Integrujesz dane z innymi narzędziami ...to **airtable database builder** prawdopodobnie zaoszczędzi Ci dziesiątki godzin miesięcznie i sporo nerwów. Sam przeszedłem tę drogę kilka lat temu i nie wyobrażam sobie powrotu do zarządzania projektami i danymi klientów w Excelu. To jak przesiadka z Nokie 3310 na smartphone'a - teoretycznie oba są telefonami, ale możliwości są nie do porównania.

Potrzebujesz pomocy z migracją z Excel do Airtable?

Pomogę Ci bezpiecznie przenieść dane, zaprojektować strukturę bazy, skonfigurować automatyzacje i przeszkolić zespół. Od analizy potrzeb przez migrację po wdrożenie i wsparcie.

Umów bezpłatną konsultację
## FAQ
### Jaka jest główna różnica między Airtable a Excel? Airtable to relacyjna baza danych z interfejsem arkusza kalkulacyjnego, Excel to arkusz kalkulacyjny. W Airtable możesz tworzyć relacje między tabelami (np. Klienci → Projekty → Faktury) bez VLOOKUP-ów, a dane automatycznie się synchronizują. Excel sprawdza się przy indywidualnej pracy i obliczeniach, Airtable przy współpracy zespołowej.
### Jak przenieść dane z Excel do Airtable krok po kroku? Przygotuj dane w Excelu (nagłówki w pierwszym wierszu, usuń puste wiersze), zaimportuj plik .xlsx do nowej bazy Airtable, dopracuj typy pól i stwórz relacje między tabelami. Na końcu dodaj widoki (Calendar, Kanban, Gallery) i skonfiguruj automatyzacje. Cały proces zajmuje od kilku godzin do kilku dni w zależności od złożoności danych.
### Kiedy lepiej zostać przy Excel zamiast migrować do Airtable? Excel wygrywa przy zaawansowanych obliczeniach finansowych i statystycznych, indywidualnej analizie danych, jednorazowych raportach oraz pracy offline bez dostępu do internetu. Jeśli pracujesz samodzielnie nad danymi bez relacji i nie potrzebujesz automatyzacji - Excel jest wystarczający.
### Co to są relacje między tabelami w Airtable i dlaczego są ważne? Relacje to połączenia między tabelami pozwalające powiązać np. artykuły z autorami jednym kliknięciem. Zamiast VLOOKUP-ów, które łamią się przy zmianach struktury, Airtable automatycznie synchronizuje powiązane dane. Kliknięcie w autora pokazuje wszystkie jego artykuły bez ręcznego filtrowania czy kopiowania danych.
### Od jakiego procesu najlepiej zacząć migrację do Airtable? Najlepiej sprawdzają się: kalendarze contentu, bazy klientów/CRM, zarządzanie projektami, listy zadań zespołowych i katalogi produktów. Wybierz jeden prosty proces, przetestuj przez tydzień na darmowym planie Airtable, sprawdź czy rozwiązuje problemy z Excelem - dopiero wtedy planuj pełną migrację.
--- # Hackathon Hacknation - Analiza Doświadczeń i Lekcja z Cyfryzacji Source: https://pawel.lipowczan.pl/blog/hackathon-hacknation-analiza-doswiadczen Published: 2025-12-12 # Hackathon Hacknation - Analiza Doświadczeń i Lekcja z Cyfryzacji Hackathon Hacknation, organizowany przez GovTech Polska, to wydarzenie o ogromnym rozmachu. Ponad 1500 uczestników, 480 tysięcy złotych w puli nagród i jeden cel: stworzyć w 24 godziny działające rozwiązanie dla wyzwań administracji publicznej. Dla naszego zespołu - który określiłbym mianem programistów juniorów - był to nie tylko rywalizacja, ale przede wszystkim poligon doświadczalny. Jestem "emerytowanym" programistą. Porzuciłem aktywne programowanie około 4 lata temu, co w świecie technologii to niemal lata świetlne - narzędzia i frameworki zmieniły się na tyle, że trzeba w pewnym sensie zaczynać od zera. Choć doświadczenie w inżynierii oprogramowania jest bardzo pomocne, to bez znajomości współczesnych języków i środowisk czasami działa się po omacku. Pozostali członkowie zespołu nigdy nie mieli zbyt wiele wspólnego z tradycyjnym programowaniem. Na co dzień wykorzystują technologie nocode wspierane przez modele językowe. ![Zespół Hacknation](/images/hacknation-team.webp) Wspólnie zderzyliśmy nasze wyobrażenia o „inteligentnych" agentach sztucznej inteligencji z twardą rzeczywistością. Oto historia o tym, jak technologia spotkała się z biurokracją, dlaczego brak walidacji może zabić najlepszy projekt i czego nauczyliśmy się o współpracy z AI pod presją czasu. ## 1. Pani Zosia i Tysiące Exceli - Analiza Problemu Nasze zadanie dotyczyło procesu budżetowania w administracji publicznej. Brzmi nudno? Może, ale skala problemu jest ogromna. Core problemem okazał się proces oparty na ręcznej wymianie setek tysięcy plików Excel. Błędy, chaos informacyjny, brak transparentności - to codzienność urzędników. Stworzyliśmy metaforę tego procesu, którą nazwaliśmy **„Pani Zosia i Tysiące Exceli”**: 1. **Start (Dół):** „Pani Zosia” w urzędzie gminy „wróży z fusów”, ręcznie wpisując dane budżetowe do Excela (np. zapotrzebowanie na nowy komputer). 2. **Eskalacja (Góra):** Plik wędruje w górę hierarchii: Urząd Miasta → Województwo → Ministerstwo Finansów. 3. **Konsolidacja:** Specjalna komórka w ministerstwie scala dane ze wszystkich plików (często ręcznie!). 4. **Decyzja i Powrót (Dół):** Limity budżetowe wracają tą samą drogą, często z arbitralnymi cięciami. W efekcie „Pani Zosia” dowiaduje się, że nie dostanie nowego komputera, ale nikt nie potrafi jej wyjaśnić dlaczego. ## 2. Rozwiązanie: Cyfrowy Budżet Postawiliśmy na proste, ale radykalne rozwiązanie: **Cyfrowy Budżet**. Zamiast przesyłać pliki, przenieśmy cały proces do chmury. Nasza koncepcja opierała się na scentralizowanej aplikacji webowej z kilkoma kluczowymi funkcjonalnościami: * **Jedno źródło prawdy:** Wszystkie pozycje budżetowe są dodawane w jednym systemie, widocznym (z odpowiednimi uprawnieniami) dla każdego szczebla. * **Transparentność i komunikacja:** Możliwość komentowania i dyskutowania nad każdą pozycją budżetową bezpośrednio w systemie, zamiast w mailach. * **Workflow akceptacji:** Uproszczony proces zatwierdzania i konsolidacji budżetu. Co ciekawe, do wyboru samego zadania również zatrudniliśmy AI. Przeanalizowaliśmy dostępne wyzwania pod kątem kompetencji naszego zespołu (głównie „nie-programistów”), aby zmaksymalizować nasze szanse. Wybór padł na budżetowanie, gdzie zrozumienie procesu biznesowego wydawało się ważniejsze niż skomplikowane algorytmy. ## 3. Atmosfera, Pot i Brak Snu ![Zakończenie Hackathonu](/images/hacknation-end.webp) Atmosfera na Hacknation była niesamowita. Wielka hala, open space, scena, ciągłe prelekcje - energia tysiąca ludzi "zajaranych technologią" udzielała się każdemu. To właśnie ten klimat pozwalał nam działać, mimo że zmęczenie narastało z każdą godziną. Praca trwała non-stop przez 24 godziny. Spaliśmy po 2-3 godziny na korytarzu lub w dedykowanej sali sypialnianej, gdzie co chwilę kogoś budził alarm telefonu. Początkowo każdy z nas rzucił się do pracy "na żywioł", tworząc własne kawałki kodu. Szybko jednak zrozumieliśmy, że to droga donikąd. Zwrot akcji nastąpił, gdy zdecydowaliśmy się skonsolidować siły wokół prototypu Justyny, który był najbardziej zaawansowany. Stał się on fundamentem naszego finalnego rozwiązania. ## 4. AI jako "Equalizer" - Technologia w Praktyce Główna teza, którą chcieliśmy sprawdzić, brzmiała: **AI to equalizer**. Narzędzie, które pozwala zespołowi z mniejszym doświadczeniem koderskim (nieprogramistom) konkurować z profesjonalnymi dev teamami. Nasz stack technologiczny: * **Frontend:** React, TypeScript * **Backend:** Supabase * **Prezentacja:** Wideo wygenerowane w HiGen Używaliśmy ciężkiej artylerii AI: * **Paweł i Kuba:** Antigravity (modele Gemini Pro / Claude 4.5). Zużyliśmy cały tygodniowy limit tokenów w kilkanaście godzin. * **Justyna:** Bolt (model Claude Code). Rekordowe zużycie **18 milionów tokenów**. ### Kontrariańskie spojrzenie na AI Czy AI napisało aplikację za nas? Nie do końca. Mimo entuzjazmu, czuliśmy lekkie zawiedzenie. * **Kod często nie działał:** AI generowało rozwiązania, które wyglądały poprawnie, ale sypały się przy uruchomieniu. * **Halucynacje:** Proponowane biblioteki nie istniały, a fragmenty logiki były "od czapy". * **Potrzeba prowadzenia za rękę:** Osiągnięcie poprawnego wyniku wymagało precyzyjnego promptowania i ciągłego korygowania kursu. * **Blokady:** Justyna napotkała błąd w filtrach _current user_, którego model nie potrafił zdiagnozować. Musieliśmy wrócić do korzeni - czytać kod i debugować ręcznie. AI to potężny mnożnik siły, ale nie magiczna różdżka. Bez umiejętności technicznych i krytycznego myślenia utknęlibyśmy w połowie drogi. ## 5. Największa Słabość - Brak Walidacji Nasz wynik końcowy to **2.15 / 5 punktów**. Nie weszliśmy do finału. Dlaczego? Technologia działała. Prezentacja była świetna. Zabrakło jednego, kluczowego elementu: **WALIDACJI**. Zespół nie miał dostępu do praktyka - urzędnika, który na co dzień pracuje z budżetem. Mentor przypisany do zadania nie był ekspertem dziedzinowym. W efekcie stworzyliśmy system, który nam wydawał się logiczny, ale mógł być kompletnie oderwany od realiów administracji („przestrzelony”). To najważniejsza lekcja: **Walidacja > Technologia**. Nawet najlepszy kod nie obroni rozwiązania, które nie odpowiada na realne potrzeby użytkownika. ## 6. Plan na Kolejny Hackathon Nauczeni doświadczeniem, przygotowaliśmy ulepszony proces na przyszłość: ### 1. Wybór zadania * Zdefiniowanie ról w zespole (mocne/słabe strony). * Scraping zadań i analiza przez agenta AI pod kątem dopasowania do zespołu. * _Wniosek:_ Ten etap mieliśmy opanowany dobrze. ### 2. Analiza biznesowa (Tutaj polegliśmy!) * Przygotowanie mapy obecnego procesu (AS-IS). * Projekt procesu docelowego (TO-BE). * Spisanie User Stories. * **PRD (Product Requirements Document):** Co budujemy i dlaczego? * **SRS (Software Requirements Specification):** Jak to zbudujemy? * Przygotowanie "skilli" dla agentów AI. ### 3. Development * Start z przygotowanym boilerplate'm (nie trać czasu na setup!). * Iteracyjny development wymagań funkcjonalnych. * Generowanie testów automatycznych przez AI. * Ciągłe Code Review. ### 4. Dokumentacja i Weryfikacja * Review pod kątem bezpieczeństwa i wydajności. * Przygotowanie dokumentacji (opcjonalnie). ## Podsumowanie i Wnioski Hacknation był dla nas bezcenną lekcją. Potwierdził, że AI pozwala robić rzeczy niemożliwe jeszcze rok temu - mały zespół w 24h stworzył działającą aplikację webową. Jednocześnie obnażył brutalną prawdę: w świecie produktu **technologia jest wtórna wobec zrozumienia problemu**. Inne zespoły, które przyszły z gotowymi komponentami i lepiej odrobiły pracę domową z analizy biznesowej, wygrały. Podejście "na żywioł" jest romantyczne, ale w starciu z przygotowaniem - przegrywa. AI to przyszłość programowania, ale to człowiek wciąż musi być pilotem, który wie, dokąd leci.

Chcesz wdrożyć AI w swojej organizacji?

Pomogę Ci znaleźć realne zastosowania AI w Twoim biznesie, uniknąć popularnych pułapek i wdrożyć rozwiązania, które przynoszą mierzalne rezultaty. Od koncepcji przez prototyp po produkcję.

Umów bezpłatną konsultację
## FAQ
### Czy AI pozwala nieprogramistom konkurować z profesjonalnymi zespołami na hackathonach? AI to equalizer - zespół bez doświadczenia koderskiego może w 24h stworzyć działającą aplikację webową. Jednak AI nie jest magiczną różdżką: kod często nie działa, modele halucynują nieistniejące biblioteki, a osiągnięcie wyniku wymaga ciągłego korygowania kursu. Bez umiejętności technicznych i krytycznego myślenia utkniesz w połowie drogi.
### Jaki jest największy błąd zespołów na hackathonach technologicznych? Brak walidacji rozwiązania z użytkownikiem końcowym. Można stworzyć technicznie działający system, który jest kompletnie oderwany od realiów - "przestrzelony". Zespoły wygrywają nie najlepszym kodem, ale lepszą analizą biznesową. Walidacja > Technologia: nawet najlepszy kod nie obroni rozwiązania, które nie odpowiada na realne potrzeby.
### Ile AI realnie pomaga przy tworzeniu projektu w 24 godziny? AI jest potężnym mnożnikiem siły - pozwala zużyć 18 milionów tokenów i wygenerować tony kodu. Ale wymaga prowadzenia za rękę: proponowane rozwiązania wyglądają poprawnie, ale sypią się przy uruchomieniu. Debugowanie nadal wymaga ręcznego czytania kodu. AI przyspiesza development, ale nie eliminuje potrzeby fundamentów technicznych.
### Jaki proces przygotowania zwiększa szanse na wygraną w hackathonie? Cztery etapy: (1) wybór zadania dopasowanego do kompetencji zespołu, (2) analiza biznesowa z mapą procesu AS-IS/TO-BE i user stories, (3) development z gotowym boilerplate i testami, (4) dokumentacja i weryfikacja. Kluczowy błąd to pomijanie etapu 2 - analiza biznesowa decyduje o sukcesie bardziej niż jakość kodu.
### Dlaczego zrozumienie problemu jest ważniejsze niż technologia na hackathonach? Technologia jest wtórna wobec zrozumienia problemu użytkownika. Zespoły z gotowymi komponentami i lepszą analizą biznesową wygrywają z zespołami, które mają lepszy kod, ale nie zwalidowały rozwiązania. Podejście "na żywioł" jest romantyczne, ale w starciu z przygotowaniem przegrywa.
--- # Kodowanie w 2025: Czy AI zbudowało moje portfolio? Case Study pawel.lipowczan.pl Source: https://pawel.lipowczan.pl/blog/kodowanie-w-2025-ai-portfolio Published: 2025-12-02 Często słyszę, że programowanie się kończy. Że wystarczy „zvibecodować” aplikację w jednym z nowych narzędzi no-code, a AI zrobi resztę. Postanowiłem to sprawdzić na żywym organizmie. Zbudowałem [pawel.lipowczan.pl](https://pawel.lipowczan.pl) - projekt, który miał być wizytówką, a stał się poligonem doświadczalnym dla współpracy na linii Doświadczony Inżynier - Agent AI. Wnioski? Jeśli myślisz, że zbudujesz profesjonalny, bezpieczny i skalowalny serwis bez wiedzy technicznej, tylko „rozmawiając” z chatbotem - jesteś w błędzie. Ale jeśli masz fundamenty inżynierskie i potraktujesz AI jako junior developera na sterydach, efekty (i koszty) mogą Cię zaskoczyć. Oto kulisy powstawania mojego portfolio w stacku React + Vite + Tailwind. ![hero](/images/hero.webp) ## 1. Fundament: Nowoczesny stack i SEO w świecie SPA Mój cel był prosty: wyjście poza ramy statycznego CV. Chciałem estetyki, wydajności i miejsca na dzielenie się wiedzą. Wybór padł na **React 18 + Vite**. Dlaczego? Bo Vite zapewnia błyskawiczny build. Jednak React to zazwyczaj Single Page Application (SPA), co bywa problematyczne dla SEO. Tutaj wchodzi inżynieria. Zastosowałem mechanizm **prerenderingu**. Mimo że pod maską działa React i Tailwind CSS, serwujemy robotom statyczne pliki HTML. Efekt? Strona jest błyskawiczna, a Google widzi ją tak, jak klasyczny dokument. Całość hostuję na **Vercel**, co okazało się strzałem w dziesiątkę. Wbudowana analityka i analiza Core Web Vitals pozwoliły mi wykręcić „zielone wyniki” niemal od razu po deployu. ![speed_insights](/images/speed_insights.webp) ![web_analytics](/images/web_analytics.webp) ## 2. Design dla nie-designera: Koniec z „wodotryskami” metodą prób i błędów Bądźmy szczerzy: nie jestem designerem. Zawsze miałem problem z doborem palety barw, ułożeniem elementów czy animacjami. W tradycyjnym modelu spędziłbym godziny na przesuwaniu pikseli w CSS. Tutaj ten problem zniknęł. Mogłem wskazać AI przykłady stron, które mi się podobają, a agent adaptował ten styl do mojego projektu. Zamiast eksperymentować z kodem CSS, wskazywałem w IDE miejsce do poprawy, a AI korygowało layout w sekundy. **Lovable vs. IDE** Eksperymentowałem z różnymi narzędziami do „vibe codingu”, m.in. z Lovable, Vercel czy Firebase Studio. Lovable dawało świetne efekty wizualne, ale finalnie zdecydowałem się na generowanie kodu bezpośrednio w IDE (Cursor). Dlaczego? Bo zależało mi na pełnej kontroli i nowoczesnym, schludnym efekcie końcowym, który jest „moim” kodem, a nie zamkniętą czarną skrzynką. ## 3. Grafika: Spójność ponad perfekcję W idealnym świecie każdy projekt w portfolio byłby opatrzony dedykowanymi zrzutami ekranu z konkretnych narzędzi i procesów. Ale przygotowanie setek takich screenów to tytaniczna praca. Zastosowałem podejście: **Done is better than perfect.** Zamiast tracić czas na robienie zrzutów czy szukanie zdjęć stockowych, postawiłem na grafikę generatywną. * Wykorzystałem **Nano Banana MCP**. * Ja dostarczam treść, agent generuje grafikę, a skrypt (konwerter) automatycznie zamienia PNG na WebP. Dzięki temu strona jest spójna wizualnie, utrzymana w klimacie tech/cyber, a ja nie muszę martwić się o "dziury" w contencie. ![og-zapier-vs-make-vs-n8n-wybor-narzedzia](/images/og-zapier-vs-make-vs-n8n-wybor-narzedzia.webp) ## 4. Rzeczywistość: Doświadczenie vs Nowe Frameworki Mam 15 lat doświadczenia w IT (.NET, Python, JS), ale frameworki takie jak React czy Vite były dla mnie nowością. Musiałem się ich poduczyć. I tu kluczowa uwaga: **Wiedza programistyczna jest niezbędna.** Dzięki doświadczeniu w projektowaniu systemów, kod generowany przez AI jest dla mnie zrozumiały. Potrafię ocenić jego poprawność, zanim trafi na produkcję. Bez tego utonąłbym w błędach. * Agenci (nawet Claude Sonnet 4.5 czy Gemini 3 Pro) potrafią się zapętlić. * Zdarzają się halucynacje nieistniejących bibliotek. Gdybym nie rozumiał fundamentów, nie byłbym w stanie „odkręcić” błędów, które AI wprowadzało przy bardziej skomplikowanej logice. ### Pedantyzm w kodzie Jestem pedantyczny, jeśli chodzi o porządek w plikach (Clean Code). Jasna struktura folderów i podział odpowiedzialności to dla mnie świętość. AI ma tendencję do wrzucania wszystkiego do jednego worka. Moja rola polegała na wymuszaniu tej struktury. Dzięki temu projekt jest łatwy w utrzymaniu i reorganizacji, a nie jest „spaghetti kodem” wyplutym przez maszynę. ## 5. Polisa ubezpieczeniowa: Testy i Code Review Agent W projekcie hobbystycznym łatwo o chaos. Aby temu zapobiec, wdrożyłem dwa poziomy zabezpieczeń: 1. **Bogate testy E2E (Playwright):** Każda zmiana jest weryfikowana przez automatyczne testy. Mam pewność, że nowa funkcja nie „rozsypała” starej. 2. **Code Review Agent w Cursor.sh:** To genialna funkcja. Agent analizuje zmiany w ostatnim commicie *przed* wysłaniem do repozytorium. Wyłapał mi sporo błędów logicznych i potencjalnych problemów, które mogłem przeoczyć. ![playwright_report](/images/playwright_report.webp) ## 6. Koszty: Ile kosztuje „darmowy” programista? To ciekawe zestawienie. Przez cały projekt przepuściłem około **60 milionów tokenów**. Rozkład wejście/wyjście to mniej więcej 80/20. Gdybym płacił cennikowo za API (np. Claude Sonnet 4.5 - $3 input / $15 output), koszt wyniósłby około **325 USD**. Realny koszt? * Plan PRO+ w Cursor.sh: **60 USD**. * Antigravity (w ramach Google Workspace): **0 USD** (wliczone w pakiet firmowy). Oszczędność jest kolosalna. Oczywiście są limity - przy intensywnej sesji zdarzało mi się zobaczyć komunikat o ich przekroczeniu. Wtedy po prostu zmieniałem model (skaczę między Gemini 3 Pro High a Claude Sonnet 4.5). Limity odnawiają się co kilka godzin, więc przy regularnej pracy nie stanowi to blokady. ![cursor_usage](/images/cursor_usage.webp) ## Podsumowanie Projekt [pawel.lipowczan.pl](https://pawel.lipowczan.pl) to dowód na to, że w 2025 roku rola programisty ewoluuje. Przestajemy być rzemieślnikami od składni, a stajemy się architektami zarządzającymi zespołem cyfrowych agentów. Możesz nie być designerem. Możesz nie znać na wylot najnowszego frameworka. Ale jeśli masz inżynierski umysł, dbałość o jakość (i testy!) oraz umiejętność orkiestracji AI - zbudujesz rzeczy, które wcześniej wymagałyby całego zespołu. Zapraszam do sprawdzenia efektów i code review! Feedback mile widziany.

Potrzebujesz wsparcia w rozwoju produktu z AI?

Pomogę Ci w wyborze technologii, projektowaniu architektury i wdrożeniu najlepszych praktyk. Od MVP przez skalowanie po optymalizację procesów deweloperskich z wykorzystaniem AI.

Umów bezpłatną konsultację
## FAQ
### Czy AI może zastąpić programistów w 2025 roku? AI to junior developer na sterydach - przyspiesza pracę 10x, ale wymaga nadzoru doświadczonego inżyniera. Agenci potrafią się zapętlić, halucynują nieistniejące biblioteki i generują "spaghetti kod" bez narzuconej struktury. Fundamenty inżynierskie są niezbędne do oceny poprawności kodu i odkręcania błędów.
### Jakie umiejętności są potrzebne programiście do efektywnej pracy z AI? Znajomość architektury systemów, clean code i podział odpowiedzialności - AI ma tendencję do wrzucania wszystkiego do jednego worka. Umiejętność czytania i oceny kodu, nawet w nieznanym frameworku. Dbałość o jakość: testy E2E i code review przed każdym commitem. Bez tych fundamentów utoniesz w błędach.
### Ile kosztuje rozwijanie projektu z pomocą AI zamiast tradycyjnego programowania? Plan PRO+ w Cursor.sh to około $60/miesiąc przy intensywnej pracy. Dla porównania: te same 60 milionów tokenów przez API kosztowałyby ~$325. Oszczędność jest kolosalna, ale są limity - przy intensywnej sesji trzeba zmieniać model lub czekać na odnowienie. Regularną pracę da się prowadzić bez blokad.
### Jak zapewnić jakość kodu generowanego przez AI? Dwa poziomy zabezpieczeń: testy E2E (Playwright) weryfikujące każdą zmianę automatycznie oraz code review agent analizujący zmiany przed commitem. Agent wyłapuje błędy logiczne i potencjalne problemy, które łatwo przeoczyć. Bez testów i review chaos w projekcie jest nieunikniony.
### Czy osoba bez doświadczenia programistycznego może zbudować profesjonalną aplikację z AI? Nie - profesjonalny, bezpieczny i skalowalny serwis wymaga wiedzy technicznej. "Vibe coding" i rozmowa z chatbotem dają efekty wizualne, ale nie pełną kontrolę nad kodem. AI doskonale wspiera doświadczonych inżynierów, ale nie zastępuje fundamentów programistycznych przy złożonych projektach.
--- # Każda firma działa nieoptymalnie - jak przestać kłamać pracownikom i zacząć naprawiać procesy? Source: https://pawel.lipowczan.pl/blog/kazda-firma-dziala-nieoptymalnie Published: 2025-12-01 Czy masz w firmie taką osobę? Filar. Kogoś, kto nigdy nie zawodzi, nawet gdy wszystko wokół się sypie. I po raz trzeci w tym miesiącu ta osoba przychodzi do Ciebie z tym samym, absurdalnym problemem - błędem w systemie, który blokuje jej pracę. Patrzysz na ekran, potem na nią, i czujesz to palące ukłucie wstydu. Mówisz: _„Spokojnie, załatwię to”_, ale w głębi duszy wiesz, że kłamiesz. Nie ze złej woli. Kłamiesz, bo technologia, która miała pomagać, robi z Ciebie kłamcę w oczach Twoich najlepszych ludzi. Jesteście skazani na skostniałe systemy, gdzie każda zmiana to projekt na miarę wyprawy na Księżyc. Jeśli ten scenariusz brzmi znajomo, nie jesteś sam. W Automation House zmapowaliśmy ponad 400 procesów i wniosek jest jeden: **każda firma działa nieoptymalnie**. Pytanie brzmi tylko: jak szybko jesteś w stanie znaleźć te miejsca i je naprawić? ## Bez mapy nie ma nawigacji Wyobraź sobie, że chcesz dotrzeć do celu w nieznanym terenie. Bez mapy błądzisz. W biznesie jest tak samo. Mapa procesu to nie tylko dokumentacja - to narzędzie nawigacyjne dla trzech grup: 1. **Biznes:** Zyskuje zrozumienie, jak _naprawdę_ działa firma (wyobrażenia zarządu często mijają się z rzeczywistością). 2. **Użytkownicy:** Otrzymują jasną instrukcję działania i szybszy onboarding. 3. **IT/Wdrożeniowcy:** Mogą precyzyjnie zaprojektować architekturę, przekazywać wiedzę na temat projektu i znajdować wąskie gardła. Dowody? Badania wskazują, że mapowanie procesów w samej tylko służbie zdrowia potrafiło skrócić czas oczekiwania pacjentów o **20-45%**. Skoro działa to w tak skomplikowanym środowisku jak szpital, zadziała też w Twojej firmie. ## Dlaczego większość map jest bezużyteczna? Wielu menedżerów próbuje mapować procesy, ale robi to źle, dobierając nieodpowiednie narzędzia: - **SIPOC (tabelki):** Świetne dla analityków, niezrozumiałe dla biznesu. Trudno z nich wyciągnąć wnioski na pierwszy rzut oka. - **BPMN (Business Process Model and Notation):** Standard korporacyjny, ale zbyt skomplikowany. Nadmiar bramek logicznych i symboli sprawia, że mapa staje się nieczytelna dla przeciętnego pracownika. - **Zwykły Flowchart:** Zbyt prosty. Pokazuje "co" się dzieje, ale często pomija "kto" i "czym" to robi. ### Złoty środek: Rozszerzony Flowchart W Automation House wypracowaliśmy własną metodę. Nasza mapa procesu musi zawierać cztery kluczowe elementy dla każdego kroku: 1. **Akcja:** Co się dzieje? 2. **Aktor:** Kto to robi? 3. **Narzędzie:** Czym to robi? (np. Excel, CRM, Slack). 4. **Tryb:** Manualny czy Automatyczny? Dzięki temu od razu widzimy, gdzie człowiek wykonuje pracę robota (kopiuj-wklej) i gdzie brakuje integracji między systemami. ## Jak znajdować "pęknięte rury"? Kiedy masz już mapę stanu obecnego (AS-IS), szukanie optymalizacji staje się proste. Szukaj miejsc, gdzie: - Występuje najwięcej błędów. - Proces trwa najdłużej. - Dane są przepisywane ręcznie (ryzyko błędu, strata czasu). - Będzie to miało największy wpływ na zespół w firmie. **Złota zasada Elona Muska:** Zanim zaczniesz cokolwiek automatyzować, zadaj sobie pytanie: _Czy ten krok w ogóle jest potrzebny?_ > „Prawdopodobnie najgorszą rzeczą jest optymalizacja czegoś, co w procesie w ogóle nie powinno się znaleźć.” Najpierw usuwaj, potem upraszczaj, a dopiero na końcu automatyzuj. ## Case Study: El Padre - Jak przyspieszyć ofertowanie o 50%? Teoria teorią, ale spójrzmy na praktykę. Agencja eventowa **El Padre** zgłosiła się do nas z problemem: tworzenie ofert (szczególnie mniejszych) było zbyt czasochłonne i mało rentowne. Wiedza o poprzednich realizacjach była rozproszona w głowach pracowników - brakowało centralnej bazy wiedzy. **Co zrobiliśmy?** 1. **Krok 1: "Ucho" procesu (Fireflies.ai)** Wdrożyliśmy narzędzie AI, które nagrywa spotkania z klientami i tworzy z nich transkrypcje. Koniec z ręcznym notowaniem i gubieniem szczegółów. 2. **Krok 2: Centralny Mózg (Airtable)** Stworzyliśmy bazę wiedzy, gdzie trafiają transkrypcje, kosztorysy i dane o projektach. 3. **Krok 3: Automatyzacja (Make & AION)** Zbudowaliśmy "asystentów AI" (w oparciu o nasze narzędzie AION), którzy realizują konkretne zadania: - **Briefing:** Generuje brief na podstawie transkrypcji rozmowy. - **Event Ideas:** Podrzuca pomysły na wydarzenie, bazując na briefie i historii agencji. - **Financial Planner:** Planuje kosztorys na podstawie briefu i historii agencji. - **Offer Generator:** Tworzy ofertę z wcześniej wygenerowanych danych kosztorysów, pomysłów i briefu. **Wyniki:** - **10-50%** szybsze przygotowywanie ofert (największy zysk przy mniejszych projektach). - **10-15%** wzrostu produktywności działu produkcji. - **30 osób** w firmie realnie wspieranych przez AI w codziennej pracy. ## Podsumowanie Technologia nie służy do tego, by komplikować życie, ale by budować operacyjną doskonałość (_Operational Excellence_). Nie musisz od razu wdrażać skomplikowanych systemów klasy ERP. Zacznij od mapy. Znajdź, gdzie Twoja firma "krwawi" czasem i energią ludzi. Jeśli chcesz przestać kłamać swoim pracownikom, że "jakoś to będzie", zacznij od zmapowania jednego procesu jeszcze w tym tygodniu.

Chcesz zmapować i zoptymalizować procesy w firmie?

Pomogę Ci znaleźć wąskie gardła w procesach, zidentyfikować miejsca do automatyzacji i wdrożyć rozwiązania, które zaoszczędzą czas Twoim pracownikom. Od mapowania przez analizę po wdrożenie.

Umów bezpłatną konsultację
## FAQ
### Dlaczego każda firma działa nieoptymalnie? Procesy narastają organicznie przez lata, wyobrażenia zarządu często mijają się z rzeczywistością, a pracownicy wykonują pracę robota (kopiuj-wklej między systemami). Z mapowania ponad 400 procesów wynika jeden wniosek: nieoptymalne miejsca są wszędzie. Pytanie brzmi tylko, jak szybko je znajdziesz i naprawisz.
### Jakie elementy powinna zawierać skuteczna mapa procesu? Cztery elementy dla każdego kroku: akcja (co się dzieje), aktor (kto to robi), narzędzie (czym - Excel, CRM, Slack) i tryb (manualny czy automatyczny). Dzięki temu od razu widać, gdzie człowiek wykonuje pracę robota i gdzie brakuje integracji między systemami.
### Jak znajdować miejsca do optymalizacji w procesach firmowych? Szukaj miejsc, gdzie: występuje najwięcej błędów, proces trwa najdłużej, dane są przepisywane ręcznie (ryzyko błędu, strata czasu), oraz gdzie zmiana będzie miała największy wpływ na zespół. Przed automatyzacją zadaj pytanie: czy ten krok w ogóle jest potrzebny?
### Dlaczego większość map procesów jest bezużyteczna? SIPOC (tabelki) jest niezrozumiałe dla biznesu, BPMN jest zbyt skomplikowane (nadmiar bramek i symboli), a zwykły flowchart zbyt prosty - pokazuje "co" bez "kto" i "czym". Złoty środek to rozszerzony flowchart z czterema elementami: akcja, aktor, narzędzie, tryb.
### Jaka jest właściwa kolejność działań przy optymalizacji procesów? Najpierw usuwaj (czy ten krok jest potrzebny?), potem upraszczaj (czy można go skrócić?), dopiero na końcu automatyzuj. Najgorszą rzeczą jest automatyzowanie czegoś, co w procesie w ogóle nie powinno się znaleźć. To złota zasada przed każdym projektem optymalizacji.
--- _Artykuł powstał na bazie prezentacji Pawła Lipowczana "Każda firma działa nieoptymalnie" wygłoszonej podczas InfoShare Katowice 2025._ --- # Zapier vs Make vs n8n - jak wybrać narzędzie automatyzacji dla Twojego zespołu? Source: https://pawel.lipowczan.pl/blog/zapier-vs-make-vs-n8n-wybor-narzedzia Published: 2025-11-17 # Zapier vs Make vs n8n - jak wybrać narzędzie automatyzacji dla Twojego zespołu? Wybór złego narzędzia automatyzacji to nie tylko stracone pieniądze - to **miesiące zmarnowanego czasu**, setki przepisanych workflow i tysiące złotych na migrację, gdy w końcu zdecydujesz się na zmianę. Widziałem to dziesiątki razy: zespoły wybierają platformę na podstawie listy funkcji, a potem utykają, bo nikt nie umie z niej korzystać. Po wdrożeniu automatyzacji dla ponad 100 klientów w **Automation House** mogę powiedzieć jedno: **nie ma uniwersalnej odpowiedzi**. Są za to konkretne kryteria, które decydują, czy dane narzędzie sprawdzi się w Twoim zespole. W tym artykule pokażę Ci **framework decyzyjny**, który pomoże wybrać między **Zapier**, **Make** i **n8n** na podstawie tego, co naprawdę ma znaczenie: kompetencji Twojego zespołu, skali operacji, budżetu i wymagań bezpieczeństwa. ## Dlaczego to nie jest tylko kwestia funkcji? Wszystkie trzy platformy robią to samo - **łączą Twoje aplikacje bez kodowania**. Ale diabeł tkwi w szczegółach: - **Zapier** ma 6000+ integracji i jest tak prosty, że Twoja mama mogłaby z niego korzystać - **Make** (dawniej Integromat) daje Ci wizualny canvas, gdzie widzisz całą logikę workflow - **n8n** to marzenie technicznego zespołu: open-source, self-hosted, nieograniczone możliwości Większość firm, które spotykam, **traci 15-25 godzin tygodniowo** na powtarzalne zadania: wprowadzanie danych, powiadomienia, aktualizacje statusów, synchronizację między platformami. Automatyzacja zabija tę stratę czasu. Ale kiedy wybierają narzędzie na podstawie feature list zamiast możliwości zespołu, uderzają w ścianę i muszą budować wszystko od nowa. ## Zapier - dla zespołów non-technical, które potrzebują rezultatów teraz ### Dla kogo? Zapier to platforma **"po prostu działa"**. Prosta konfiguracja trigger-action, którą non-technical zespoły mogą wdrożyć w kilka minut. **Najlepszy dla:** - Działów RevOps łączących HubSpot/Salesforce/Slack - Marketing automation i lead routing - IT workflows (onboarding, ticketing) - Startupów bez technicznego CTO ### Zalety ✅ **6000+ integracji** - jeśli aplikacja istnieje, Zapier ją wspiera ✅ **Zero krzywej uczenia** - non-technical zespoły zaczynają w 5 minut ✅ **Świetna dokumentacja** - template library, community, video guides ✅ **Natychmiastowe rezultaty** - pierwsze workflow w 10 minut ✅ **Największa społeczność** - każdy problem ma rozwiązanie na forum ### Wady ❌ **Limity zadań rosną jak rakieta** - 5-krokowy Zap × 100 uruchomień = **500 zadań** ❌ **Cena eskaluje szybko** - darmowe 100 zadań znika w mgnieniu oka ❌ **Ograniczona logika warunkowa** - trudno budować złożone decyzje ❌ **Debugowanie to koszmar** - kiedy coś nie działa, ciężko znaleźć przyczynę ❌ **Vendor lock-in** - migracja do innej platformy = przepisywanie od zera ### Przykłady zastosowań **Lead routing:** ```text Formularz kontaktowy → Zapier → - Dodaj lead do HubSpot - Wyślij powiadomienie na Slack - Stwórz zadanie w Asana - Wyślij email powitalny ``` **Onboarding pracownika:** ```text Nowy rekord w BambooHR → Zapier → - Utwórz konto w Google Workspace - Dodaj do Slack channels - Wyślij welcome email z checklistą - Stwórz zadania dla managera ``` ### Koszty - **Free:** 100 zadań/miesiąc - **Starter:** $19.99 (750 zadań) - **Professional:** $49 (2,000 zadań) - **Team:** $299 (50,000 zadań) **⚠️ Uwaga:** Każdy krok w Zapie to osobne zadanie! 5-krokowy Zap uruchomiony 100 razy = 500 zadań. ### Kiedy wybierać Zapier? ✅ Twój zespół jest non-technical (marketing, sales, ops) ✅ Potrzebujesz rezultatów natychmiast, bez szkoleń ✅ Łączysz niszowe aplikacje (mają największą liczbę integracji) ✅ Prowadzisz piloty i testy koncepcyjne ✅ Skalowanie nie jest Twoim priorytetem (< 5,000 zadań/miesiąc) ## Make - dla wizualnych myślicieli z ambicjami ### Dla kogo? Make to platforma dla zespołów, które **myślą wizualnie** i potrzebują więcej mocy niż Zapier, ale bez technicznej złożoności n8n. **Najlepszy dla:** - Agencji kreatywnych z content pipelines - Marketing automation z personalizacją - Zespołów z power userami - Procesów wymagających złożonej logiki "if-this-then-that" ### Zalety ✅ **Visual workflow builder** - widzisz cały proces na canvas ✅ **Lepszy stosunek ceny do możliwości** - 10x więcej operacji za te same pieniądze ✅ **Zaawansowana logika** - routers, filters, iterators, error handlers ✅ **Transparentne debugowanie** - każdy krok pokazuje dane wejściowe/wyjściowe ✅ **Operacje ≠ kroki** - każdy krok to 1 operacja (nie mnoży się jak w Zapier) ### Wady ❌ **Starsza krzywa uczenia** - visual builder wymaga oswojenia się ❌ **Mniej integracji** - 1800+ aplikacji (vs 6000+ w Zapier) ❌ **Dokumentacja nierówna** - niektóre moduły słabo udokumentowane ❌ **Interfejs może przytłaczać** - początkowo chaos na canvas ### Przykłady zastosowań **Content pipeline z kategorizacją:** ```text Webhook → Make → ├─ Jeśli typ = "blog post" │ └─ Dodaj do WordPress + notify writers ├─ Jeśli typ = "social media" │ └─ Schedule w Buffer + notify social team └─ Jeśli typ = "newsletter" └─ Dodaj do Mailchimp + notify subscribers ``` **Marketing campaign z personalizacją:** ```text Nowy subscriber → Make → ├─ Pobierz dane z CRM ├─ Router według segmentu: │ ├─ B2B → Email sequence A │ ├─ B2C → Email sequence B │ └─ Enterprise → Notify sales team └─ Dodaj do odpowiedniej listy remarketing ``` ### Koszty - **Free:** 1,000 operacji/miesiąc - **Core:** $9 (10,000 operacji) - **Pro:** $16 (10,000 operacji + premium apps) - **Teams:** $29 (10,000 operacji + team features) **🎯 Kluczowa różnica:** W Make każdy krok = 1 operacja (nie mnoży się!). 10-krokowy workflow × 1000 uruchomień = 10,000 operacji. ### Kiedy wybierać Make? ✅ Chcesz więcej mocy niż Zapier bez technicznej złożoności n8n ✅ Twój zespół myśli wizualnie i lubi "widzieć" logikę ✅ Workflow mają wiele rozgałęzień i warunków ✅ Budujesz dla klientów i musisz pokazywać logikę ✅ Szukasz najlepszego stosunku ceny do możliwości ## n8n - dla technical teams z wymaganiami ### Dla kogo? n8n to **open-source powerhouse** dla zespołów technicznych, które chcą pełnej kontroli nad automatyzacją. **Najlepszy dla:** - Zespołów z developerami/DevOps - Branż regulowanych (healthcare, finance, legal) - High-volume automation (50K+ zadań/miesiąc) - Potrzeby integracji custom API - Agencji budujących produkty automatyzacji dla klientów ### Zalety ✅ **Open-source (MIT license)** - pełny dostęp do kodu źródłowego ✅ **Self-hosted = zero kosztów subskrypcji** - płacisz tylko za infrastrukturę ✅ **Nieograniczone workflow steps** - brak limitów kroków ✅ **Custom code nodes** - JavaScript w każdym kroku ✅ **Pełna kontrola nad danymi** - dla compliance (HIPAA, GDPR, SOC2) ✅ **API-first approach** - łatwa integracja z custom systems ✅ **Świetne dla AI agents** - zaawansowane workflow z LLM ### Wady ❌ **Wymaga DevOps skills** - Docker, databases, SSL, backups, monitoring ❌ **Self-hosting = maintenance overhead** - aktualizacje, security patches ❌ **Mniejsza społeczność** - mniej template'ów i przykładów ❌ **Cloud hosting droższy** - niż Make (jeśli nie self-hostujesz) ❌ **Odpowiedzialność za security** - sam musisz dbać o bezpieczeństwo ### Przykłady zastosowań **Healthcare data pipeline (HIPAA compliant):** ```text Patient intake form → n8n (self-hosted) → ├─ Encrypt PHI data ├─ Store in compliant database ├─ Notify medical staff (secure channel) └─ Log audit trail ``` **AI agent workflow:** ```text User query → n8n → ├─ Pre-process with custom code ├─ Route do odpowiedniego LLM (OpenAI/Claude/Local) ├─ Post-process response ├─ Store w vector database └─ Return formatted result ``` **Multi-tenant automation product:** ```text Client webhook → n8n → ├─ Identify tenant ├─ Load tenant-specific config ├─ Execute custom workflow ├─ Bill based on usage └─ Store metrics per tenant ``` ### Koszty **Self-hosted:** - **Software:** $0 (MIT license) - **Infrastructure:** $10-50/miesiąc (VPS: DigitalOcean, Hetzner, AWS) - **DevOps time:** 5-10h/miesiąc (setup, maintenance) **n8n Cloud:** - **Starter:** $20 (2,500 workflow executions) - **Pro:** $50 (10,000 executions) - **Enterprise:** Custom pricing ### Kiedy wybierać n8n? ✅ Masz developera lub DevOps osobę w zespole ✅ Data privacy jest krytyczna (healthcare, finance, legal) ✅ Skalujesz powyżej 50K zadań/miesiąc ✅ Potrzebujesz custom code lub niestandardowych API ✅ Budujesz produkty automatyzacji dla wielu klientów (multi-tenant) ✅ Wymagania compliance (HIPAA, GDPR on-premises) ## Kod z AI agentami - opcja, której nikt nie bierze pod uwagę (a powinna) ### Dla kogo? To zmienił rok 2026. Agenty AI jak Claude Code, Cursor czy GitHub Copilot zrobiły z pisania kodu to, co Zapier zrobił z integracjami: **obniżyły barierę wejścia do zera**. **Najlepszy dla:** - Każdego kto ma developera (nawet juniora) z dostępem do AI agenta - Firm które chcą pełnej kontroli bez vendor lock-in - Use-cases wymagających niestandardowej logiki i wielu integracji - Każdego kto buduje coś specyficznego dla swojego biznesu (micro-tools) ### Jak to działa w 2026? Dawna kalkulacja: własny kod = tygodnie pracy, high cost, trudne utrzymanie. **Nowa kalkulacja**: masz API dokumentację? Dajesz ją agentowi AI. Integracja gotowa w godziny, nie tygodnie. Przykład: potrzebujesz webhook, który bierze dane z Notion, wzbogaca je przez zewnętrzne API, filtruje według 5 kryteriów i zapisuje do Airtable + wysyła Slack? W Make: 45 minut konfiguracji. W kodzie z Claude Code: 2 godziny i masz deployment. Ale kod jest Twój. Nie masz limitu na uruchomienia. Nie płacisz $50/miesiąc na zawsze. ### Zalety ✅ **Zero vendor lock-in** - kod działa wszędzie, migrujesz gdzie chcesz ✅ **Każde API bez oczekiwania** - daj dokumentację agentowi, ma integrację w godziny ✅ **Brak limitów** - zero artificial limits na kroki, uruchomienia, dane ✅ **Najtańszy przy skali** - hosting $5-20/miesiąc vs setki dolarów subskrypcji ✅ **Pełna debugowalność** - żadnych czarnych skrzynek, każdy krok w logach ✅ **Composable** - jak klocki Lego: każdy micro-tool robi jedno i robi to dobrze ✅ **Najlepsza elastyczność** - zmiana logiki to zmiana kodu, nie walka z UI platformy ### Wady ❌ **Wymaga developera** - junior + AI agent to minimum, ale to musi być człowiek z tech background ❌ **Maintenance** - kod trzeba utrzymywać (choć AI pomaga też z tym) ❌ **Czas setupu** - pierwsze wdrożenie wolniejsze niż "klik-klik w Zapierze" ❌ **Infrastructure** - hosting, deployment, monitoring (ale narzędzia jak Railway/Fly.io minimalizują overhead) ### Przykład: micro-tool zamiast platformy **Zadanie**: Przetworzyć 500 leadów dziennie z 3 źródeł, zdeduplikować, wzbogacić danymi z Clearbit, zapisać do CRM i powiadomić sales. **W Zapierze**: $300+/miesiąc, 5-krokowy Zap × 500 = 2500 zadań/dzień **W Make**: $200+/miesiąc, złożony scenario **W kodzie + Claude Code**: 1-2 dni budowy, $10/miesiąc na infrastrukturze, nieograniczone ```python # To co AI agent napisze dla Ciebie w godziny: # - Webhook przyjmujący leady z 3 źródeł # - Deduplication logic # - Clearbit enrichment # - CRM write # - Slack notification # Zero platformy. Zero limitu. Zero vendor lock-in. ``` ### Koszty - **Infrastruktura**: $5-20/miesiąc (Railway, Fly.io, Render) - **Czas developera**: 2-8h jednorazowo (z AI agentem zamiast 2-4 tygodni) - **Maintenance**: 1-2h/miesiąc (z pomocą AI) - **Total**: ~$20/miesiąc + jednorazowy koszt ### Kiedy wybierać Kod z AI agentami? ✅ Masz developera (junior/mid) z dostępem do Claude Code / Cursor ✅ Chcesz zero vendor lock-in ✅ Logika jest specyficzna dla Twojego biznesu ✅ Skala >20K operacji/miesiąc (gdzie subskrypcje bolą) ✅ Chcesz budować micro-tools, nie wdrażać monolityczną platformę ✅ Integracje z API które nie mają oficjalnych konektorów w Zapier/Make ## Framework decyzyjny - jak właściwie wybrać? Zamiast zgadywać, użyj tego prostego frameworka: ### Pytanie 1: Jakie kompetencje ma Twój zespół? - **Non-technical** (marketing, sales, operations) → **Zapier** - **Power users**, wizualni myśliciele → **Make** - **Developerzy**, DevOps, technical team → **n8n** ### Pytanie 2: Jaka jest skala operacji? - **< 5,000 zadań/miesiąc** → **Zapier** lub **Make** - **5,000 - 50,000 zadań/miesiąc** → **Make** - **> 50,000 zadań/miesiąc** → **n8n** (self-hosted) ### Pytanie 3: Jaki jest budżet? - **Minimalny budżet, quick wins** → **Zapier Free/Starter** - **Najlepsza wartość za pieniądze** → **Make** - **Long-term, high-volume** → **n8n self-hosted** ### Pytanie 4: Jakie wymagania bezpieczeństwa? - **Standard SaaS security** → **Zapier** / **Make** - **Data residency, compliance** → **n8n self-hosted** - **HIPAA, GDPR, SOC2 on-premises** → **n8n self-hosted** ### Pytanie 5: Jaka złożoność procesów? - **Proste trigger-action** (A → B → C) → **Zapier** - **Multi-step z warunkami** (if-else, routers) → **Make** - **Complex logic + custom code** → **n8n** ### Pytanie 6: Czy masz developera z dostępem do AI agenta? - **Tak** → rozważ Kod z AI jako pierwszą opcję (zero vendor lock-in, najniższy koszt przy skali) - **Nie** → wróć do pytań 1-5 (Zapier/Make/n8n) ## Typowe błędy przy wyborze (i jak ich uniknąć) ### Błąd 1: Wybór n8n bez technical resources **Co się dzieje:** - Deploy na VPS, wszystko działa - Po tygodniu: problem z SSL certificate - Po miesiącu: baza danych pełna, backup nie działa - Po 3 miesiącach: security vulnerability, brak aktualizacji **Rozwiązanie:** ✅ Zatrudnij DevOps konsultanta (5-10h/miesiąc) ✅ Użyj managed n8n hosting (dużo droższe, ale bez headache) ✅ Lub... wybierz Make zamiast n8n ### Błąd 2: Start na Zapier i utknięcie na limicie zadań **Co się dzieje:** - Start z Zapier Free (100 zadań) - Po tygodniu: upgrade do Starter ($20, 750 zadań) - Po miesiącu: upgrade do Professional ($50, 2000 zadań) - Po kwartale: $300/miesiąc, a workflow są proste **Dlaczego?** 5-step Zap × 100 uruchomień = 500 zadań! **Rozwiązanie:** ✅ Jeśli widzisz, że skalujesz powyżej 5K zadań, **od razu idź w Make** ✅ Prototypuj w Zapier, production w Make ✅ Migracja Zapier → Make to przepisanie od zera (planuj z wyprzedzeniem) ### Błąd 3: Wybór na podstawie feature list zamiast team fit **Co się dzieje:** - CTO wybiera n8n bo "jest open-source i ma wszystkie features" - Zespół marketingu nie umie z niego korzystać - Developerzy nie mają czasu budować workflow - Rezultat: 0 wdrożonych automatyzacji po 3 miesiącach **Rozwiązanie:** ✅ **Najlepsze narzędzie to to, którego będzie używać zespół** ✅ Prostota > funkcjonalność (jeśli nikt nie umie z niej korzystać) ✅ Zacznij od pilot project z zespołem, który będzie używać narzędzia ### Błąd 4: Nie uwzględnianie Total Cost of Ownership (TCO) **Zapier TCO:** - Subscription: $50-300/miesiąc - Team time: 2h/miesiąc (maintenance) - **Total: $50-300/miesiąc** **Make TCO:** - Subscription: $16-50/miesiąc - Learning curve: 10h (jednorazowo) - Team time: 3h/miesiąc (maintenance) - **Total: $16-50/miesiąc** **n8n TCO (self-hosted):** - Infrastructure: $20-50/miesiąc - DevOps time: 10h/miesiąc × $50/h = $500 - **Total: $520-550/miesiąc** **n8n TCO (cloud):** - Subscription: $50-200/miesiąc - Team time: 3h/miesiąc - **Total: $50-200/miesiąc** **Kod + AI agent TCO:** - Infrastructure: $10-20/miesiąc (Railway, Fly.io, Render) - Dev time: 2-8h jednorazowo (z AI agentem) - Maintenance: 1-2h/miesiąc - **Total: ~$20/miesiąc + jednorazowy koszt budowy** | Narzędzie | Miesięczny koszt | Jednorazowy koszt | |--------------------|------------------|-------------------| | Zapier | $50-300 | minimal | | Make | $16-50 | low | | n8n self-hosted | $520-550 | high | | Kod + AI agent | $10-20 | medium (1x) | **Wniosek:** n8n self-hosted ma sens tylko przy **high-volume** (>50K zadań) lub **compliance requirements**. Kod + AI agent wygrywa gdy masz developera i chcesz pełnej kontroli bez rosnących subskrypcji. ## Strategie migracji między platformami ### Z Zapier do Make **Kiedy?** Koszty Zapier > $100/miesiąc, a workflow są średniej złożoności. **Jak?** 1. Zidentyfikuj najprostsze Zapy (3-5 kroków) 2. Przepisz je w Make (visual canvas ułatwia optymalizację) 3. Testuj równolegle przez tydzień 4. Wyłącz Zapy dopiero po weryfikacji 5. Stopniowo migruj bardziej złożone workflow **Czasochłonność:** 2-4h na workflow ### Z Make do n8n **Kiedy?** Koszty Make > $200/miesiąc lub compliance requirements. **Jak?** 1. Deploy n8n na managed hosting (Railway, Render) 2. Eksportuj workflow z Make jako JSON (częściowo kompatybilne) 3. Migruj najpierw non-critical workflows 4. Testuj dokładnie (różnice w node'ach) 5. Stopniowa migracja produkcyjnych workflow **Czasochłonność:** 5-10h na workflow (przepisywanie niemal od zera) ### Multi-platform approach Nie musisz wybierać tylko jednej platformy! **Strategia:** - **Zapier** - quick wins, prototypy, testy koncepcyjne - **Make** - production workflows, standardy zespołu - **n8n** - high-volume, sensitive data, complex logic **Przykład:** - Marketing używa Zapier (lead routing, proste integracje) - Product team używa Make (onboarding, notifications) - Engineering team używa n8n (data pipelines, AI agents) ## Case studies - real world scenarios ### Case Study 1: Startup marketingowy wybrał Zapier **Zespół:** 3 osoby (CEO, marketer, designer) **Problem:** Manualne lead routing z 5 źródeł **Rozwiązanie:** 3 proste Zapy **Workflow:** ```text 1. Formularz → HubSpot + Slack 2. LinkedIn Lead Gen → HubSpot + Email 3. Chatbot → HubSpot + Asana task ``` **Rezultaty:** - ✅ Wdrożenie: 2 godziny - ✅ ROI: pierwszy dzień (zaoszczędzili 5h/tydzień) - ✅ Koszty: $50/miesiąc (Professional plan) - ✅ Zadowolenie: 10/10 **Dlaczego Zapier?** Non-technical zespół, proste workflow, natychmiastowy rezultat. ### Case Study 2: Agencja kreatywna wybrała Make **Zespół:** 15 osób (designers, copywriters, project managers) **Problem:** Content chaos - 5 źródeł treści, 10 kanałów publikacji **Rozwiązanie:** 25 złożonych workflow z kategorizacją **Workflow (przykład):** ```text Content submission → Make → ├─ Classify content type (AI) ├─ Router według typu: │ ├─ Blog → WordPress + notify writers │ ├─ Social → Buffer (multi-channel) + notify social │ ├─ Newsletter → Mailchimp + notify subscribers │ └─ Client → Dropbox + notify client success ├─ Update project status (Asana) └─ Log metrics (Google Sheets) ``` **Rezultaty:** - ✅ Zaoszczędzili: 20h/tydzień - ✅ Koszty: $150/miesiąc (vs $800 w Zapier) - ✅ Workflow: 25 aktywnych, średnio 12 kroków każdy - ✅ Kompleksowość: niemożliwe do osiągnięcia w Zapier **Dlaczego Make?** Power users, wizualna logika, najlepsza wartość za pieniądze. ### Case Study 3: Software house wybrał n8n **Zespół:** 30 osób (15 developerów, 10 product, 5 ops) **Problem:** 50K+ zadań/miesiąc, wymagania GDPR **Rozwiązanie:** n8n self-hosted na AWS **Workflow (przykłady):** ```text 1. User registration → Encrypt PII → Store EU database → Email 2. Payment webhook → Process → Update CRM → Generate invoice 3. Support ticket → Classify (AI) → Route → Notify → Track SLA 4. CI/CD webhook → Test → Deploy → Notify → Update docs ``` **Rezultaty:** - ✅ Volume: 50,000+ workflow executions/miesiąc - ✅ Koszty: $30/miesiąc (AWS EC2 t3.medium) - ✅ vs Make: $400+/miesiąc (na tym volume) - ✅ vs Zapier: $1,200+/miesiąc - ✅ Compliance: GDPR-compliant (EU-hosted) **Dlaczego n8n?** Technical team, high-volume, compliance requirements, ROI po 2 miesiącach. ### Case Study 4: Software startup wybrał Kod + AI **Zespół:** 2 developerów + Claude Code **Problem:** 30K leadów/miesiąc z 5 źródeł, złożona logika deduplikacji i wzbogacania danych **Rozwiązanie:** Python micro-service + webhooks, deployment na Railway **Workflow:** ```text Webhook (5 źródeł) → ├─ Deduplication logic (custom) ├─ Clearbit enrichment ├─ Filtracja według kryteriów biznesowych ├─ Zapis do CRM └─ Slack notification ``` **Rezultaty:** - ✅ Build time: 3 dni (vs szacowane 3 tygodnie bez AI) - ✅ Koszty: $15/miesiąc (vs $400 w Make przy tej skali) - ✅ Vendor lock-in: $0 (zero) - ✅ Custom logic: nieograniczona, zmiana = zmiana linii kodu **Dlaczego Kod?** Developer + Claude Code = zwinność no-code + moc kodu. Przy 30K leadów/miesiąc Make kosztowałby $200-400/miesiąc. Kod: $15/miesiąc i nieograniczone operacje. ## Przyszłość automatyzacji no-code ### Trendy, które obserwuję **1. AI-driven automation** - ChatGPT/Claude nodes w każdej platformie - Inteligentna kategorizacja i routing - Generowanie treści w workflow **2. Conversational workflow creation** - "Stwórz workflow, który robi X" → gotowe - Citizen developers vs technical teams - Democratyzacja automatyzacji **3. Konsolidacja platform** - All-in-one (automation + data + AI) - Kestra, Temporal, Prefect - nowa generacja **4. Regulatory compliance automation** - GDPR, HIPAA, SOC2 out-of-the-box - Automated audit trails - Self-hosted renaissance ### Moje rekomendacje na 2026 **Dla startupów:** - Start simple: **Zapier** - Scale smart: **Make** gdy przekroczysz 5K zadań - Masz developera + AI: **Kod** od razu - zero vendor lock-in - Go technical: **n8n** jeśli compliance lub high-volume **Dla agencji:** - Default choice: **Make** (best value) - Client work: Zapier dla prostych, Make dla złożonych - Product building: **n8n** dla multi-tenant SaaS - Wewnętrzne narzędzia: **Kod + AI** dla custom micro-tools **Dla enterprise:** - Departmental: **Zapier**/Make dla poszczególnych działów - Central automation: **n8n** self-hosted dla IT - Tech team + AI: **Kod** dla specyficznych narzędzi i integracji - Governance: Multi-platform approach z central oversight ## Podsumowanie - Quick Decision Guide ### Chcesz najprostszy start? → **Zapier** - Non-technical zespół - < 5,000 zadań/miesiąc - Proste integracje - Rezultaty w 10 minut ### Chcesz najlepszą wartość? → **Make** - Power users w zespole - 5,000 - 50,000 zadań/miesiąc - Złożona logika warunkowa - 10x więcej za te same pieniądze ### Chcesz maksymalną kontrolę? → **n8n** - Technical zespół - > 50,000 zadań/miesiąc - Compliance requirements - Custom integrations ### Chcesz maksymalną elastyczność i zero limitów? → **Kod + AI agent** - Masz developera z dostępem do Claude Code / Cursor - Specyficzna logika biznesowa - Skala > 20K operacji/miesiąc - Zero vendor lock-in, zero artificial limits ### Nie masz pewności? → **Multi-platform approach** - Zapier dla prototypów - Make dla production - n8n dla specific use cases **Pamiętaj:** Najlepsze narzędzie to to, którego będzie używać Twój zespół. Dopasuj platformę do kompetencji zespołu, nie na odwrót.

Potrzebujesz pomocy w wyborze narzędzia do automatyzacji?

Pomogę Ci przeanalizować potrzeby Twojej firmy, wybrać odpowiednią platformę (Zapier, Make lub n8n) i wdrożyć pierwsze workflow. Od audytu procesów przez wybór narzędzi po szkolenia zespołu.

Umów bezpłatną konsultację
## FAQ
### Jak wybrać między Zapier, Make i n8n dla swojego zespołu? Wybór zależy od trzech czynników: kompetencji zespołu, skali operacji i budżetu. Zapier dla zespołów non-technical (<5K zadań/miesiąc), Make dla power userów szukających najlepszej wartości (5-50K zadań), n8n dla technical teams z wymaganiami compliance (>50K zadań).
### Która platforma automatyzacji oferuje najlepszy stosunek ceny do możliwości? n8n oferuje najlepszą wartość przy złożonych workflow - 1 kredyt to uruchomienie całego workflow niezależnie od liczby modułów. W Make każdy moduł zużywa osobny kredyt. Dla prostych workflow (3-5 kroków) Make może być bardziej opłacalne dzięki niższej cenie pojedynczego kredytu, ale przy rozbudowanych procesach n8n wygrywa ekonomicznie.
### Czy n8n self-hosted rzeczywiście jest tańsze od Zapier i Make? Tylko przy bardzo dużej skali (>50K zadań/miesiąc) lub wymaganiach compliance. Koszt infrastruktury to $20-50/miesiąc, ale dochodzi 10h/miesiąc pracy DevOps ($500+). Make za $50/miesiąc obsługuje większość przypadków bez maintenance overhead.
### Kiedy warto używać kilku platform automatyzacji jednocześnie zamiast jednej? Multi-platform approach sprawdza się gdy różne działy mają różne potrzeby. Typowa strategia: Zapier dla prototypów i szybkich testów, Make dla production workflows, n8n dla high-volume lub wrażliwych danych. Eliminuje to kompromisy wynikające z wyboru jednego narzędzia.
### Dlaczego migracja z Zapier do Make wymaga przepisania workflow od zera? Platformy używają różnych modeli danych i struktur workflow - nie ma bezpośredniej kompatybilności. Zapier liczy każdy krok jako osobne zadanie, Make traktuje cały workflow jako jedną operację. Migracja wymaga 2-4h na workflow, ale długoterminowo zwraca się przez 5-10x niższe koszty operacyjne.
--- # Jak agencja eventowa El Padre przyspieszyła tworzenie ofert nawet o 50% dzięki AI Source: https://pawel.lipowczan.pl/blog/el-padre-automatyzacja-ofert-ai Published: 2025-11-16 # Jak agencja eventowa El Padre przyspieszyła tworzenie ofert nawet o 50% dzięki AI W konkurencyjnym świecie eventów, gdzie każda minuta ma znaczenie, agencja eventowa **El Padre** musiała zmierzyć się z wyzwaniem, które dotyka wielu firm w branży: jak zwiększyć liczbę składanych ofert bez utraty ich wysokiej jakości? Odpowiedzią okazało się wdrożenie platformy **AION** z zaawansowanym wsparciem AI. Rezultaty? **10-50% szybsze przygotowywanie ofert**, **10-15% większa produktywność działu produkcji** i **25-30 osób wspieranych przez AI** w codziennej pracy. W tym case study pokażę Ci krok po kroku, jak przeprowadziliśmy wdrożenie i jakie konkretne korzyści biznesowe przyniosło to rozwiązanie. ## Problem biznesowy: Kiedy jakość spotyka się z presją czasu El Padre to renomowana agencja eventowa, która przygotowuje kompleksowe oferty dla swoich klientów - od koncepcji kreatywnej, przez wizualizacje, aż po szczegółowe kosztorysy i harmonogramy. Każda oferta to autorski projekt, który wymaga zaangażowania dwóch kluczowych działów: ### Czas to pieniądz - dosłownie Przed wdrożeniem AION, przygotowanie jednej oferty zajmowało: - **Działowi kreatywnemu:** 6-10 godzin (research, koncepcja, wizualizacje) - **Działowi produkcji:** 4-6 godzin (wycena, kosztorysowanie, logistyka) - **Łącznie:** nawet **16 godzin na jedną ofertę** Przy mniejszych projektach (50-100 tys. zł) taki nakład pracy był **nierentowny** - ale bez profesjonalnej oferty trudno było wygrać przetarg. ### Kluczowe wyzwania 1. **Przeciążenie zespołu kreatywnego** - zbyt wiele projektów, zbyt mało czasu na każdy z nich 2. **Czasochłonne prace działów kreacji i produkcji** - manualne kosztorysowanie, szukanie poprzednich ofert, tworzenie wizualizacji 3. **Brak centralizacji wiedzy** - transkrypcje spotkań, kosztorysy i archiwalne oferty były rozproszone w różnych folderach OneDrive 4. **Niewspółmierny koszt przygotowania ofert** - szczególnie przy mniejszych projektach, gdzie marża nie uzasadniała nakładu pracy Agencja stanęła przed dylematem: albo zwiększyć zespół (co generuje koszty), albo znaleźć sposób na przyspieszenie procesów bez utraty jakości. ## Rozwiązanie: AION jako centralny hub procesów ofertowania Zdecydowaliśmy się na wdrożenie platformy **AION** - systemu zarządzania procesami wspieranymi przez AI. Kluczem do sukcesu było nie tylko wprowadzenie narzędzi AI, ale przede wszystkim **integracja i centralizacja danych** oraz **dostosowanie workflow do realnych potrzeb zespołu**. ### Stack technologiczny - **AION** - platforma do zarządzania procesami wspieranymi AI - **OneDrive** - integracja z istniejącym systemem plików (automatyczna synchronizacja) - **Narzędzia AI do transkrypcji** - automatyczne nagrywanie i przetwarzanie spotkań - **Asystenci AI** - punktowe wsparcie dla konkretnych etapów ofertowania - **System przeszukiwania bazy wiedzy** - szybkie odnajdywanie informacji z poprzednich projektów ## Implementacja: Jak to zrobiliśmy w 6 tygodni Proces wdrożenia podzieliliśmy na trzy główne etapy, które zrealizowaliśmy w ciągu około **6 tygodni**: ### Krok 1: Integracja i centralizacja danych (Tydzień 1-2) Pierwszym krokiem było uporządkowanie chaosu informacyjnego. Zautomatyzowaliśmy synchronizację plików z OneDrive do platformy AION, tworząc jedno, centralne miejsce na: - **Transkrypcje spotkań z klientami** - każda rozmowa automatycznie nagrywana i przetwarzana - **Wiedzę o wszystkich projektach** - historia współpracy, preferencje, uwagi - **Kosztorysy i wyceny** - baza cenowa i szablony kalkulacji - **Archiwalne oferty** - jako baza wzorców i punkty odniesienia dla nowych projektów Dzięki temu zespół zyskał **natychmiastowy dostęp** do całej wiedzy organizacyjnej - bez przeszukiwania dziesiątek folderów. ### Krok 2: Wdrożenie inteligentnych narzędzi (Tydzień 3-4) W drugim etapie uruchomiliśmy narzędzia AI, które miały wspierać konkretne zadania: - **Automatyczne nagrywanie i transkrypcja spotkań** - każde spotkanie z klientem było nagrywane (za zgodą), a następnie przetwarzane na transkrypcję i strukturalne dane - **Przetwarzanie transkrypcji na dane strukturalne** - AI wyciągało kluczowe informacje: budżet, preferencje, wymagania, terminy - **Generator wizualizacji** - AI pomagało w szybkim tworzeniu wstępnych koncepcji wizualnych dla ofert eventowych - **System przeszukiwania bazy wiedzy** - możliwość filtrowania po spotkaniach, projektach, klientach i szybkiego odnajdywania podobnych przypadków ### Krok 3: Implementacja workflow i wsparcie (Tydzień 5-6) Ostatni etap to dostosowanie systemu do codziennej pracy i przeszkolenie zespołu: - **Przygotowanie punktowych asystentów AI** - każdy etap procesu ofertowania (research, koncepcja, wycena, finalizacja) otrzymał dedykowanego asystenta AI - **Wsparcie techniczne na etapie wdrożenia** - daily standups, rozwiązywanie problemów na bieżąco - **Przeszkolenie zespołu** - **25-30 osób** przeszło przez warsztaty i onboarding z AION - **Testy pilotażowe i optymalizacja** - dostrajanie workflow na podstawie feedbacku zespołu ## Rezultaty: Liczby mówią same za siebie Po 6 tygodniach wdrożenia El Padre zaczęła widzieć pierwsze efekty, które z czasem tylko rosły: ### Kluczowe metryki - ✅ **10-50% szybsze przygotowywanie ofert** - w zależności od złożoności projektu (proste oferty nawet 50% szybciej, złożone ~10-20%) - ✅ **10-15% większa produktywność działu produkcji** - dzięki automatyzacji wycen i kosztorysów - ✅ **25-30 osób wspieranych przez AI** - praktycznie cały zespół korzysta z AION w codziennej pracy - ✅ **Drastyczny wzrost liczby składanych ofert miesięcznie** - bez zwiększania zespołu ### ROI: Oszczędności czasowe Liczby są imponujące: - **Średnio 5-8 godzin zaoszczędzonego czasu na jedną ofertę** - Przy **15 ofertach miesięcznie** = **75-120 godzin oszczędności** - To równowartość **2-3 pełnoetatowych pracowników** ### Korzyści biznesowe **Wzrost efektywności:** - Dział produkcji może obsłużyć więcej projektów bez dodatkowych zatrudnień - Zespół kreatywny ma więcej czasu na innowacyjne koncepcje - Możliwość składania ofert na mniejsze projekty (wcześniej nierentowne) **Jakość i spójność:** - Wszystkie oferty utrzymują wysoki standard dzięki dostępowi do bazy wiedzy - Mniejsze ryzyko błędów w wycenach (automatyzacja kosztorysów) - Szybszy dostęp do archiwalnych projektów jako wzorców **Wizerunek i konkurencyjność:** - Wzmocnienie wizerunku jako innowacyjnej, nowoczesnej agencji - Szybszy czas odpowiedzi na zapytania ofertowe (przewaga konkurencyjna) - Możliwość obsługi większej liczby klientów **Dodatkowe korzyści:** - Centralizacja wiedzy - łatwiejsze wdrażanie nowych pracowników - Lepsza dokumentacja projektów - wszystko w jednym miejscu - Skalowalność - rozwiązanie rośnie wraz z firmą ## Głos klienta: Co mówi El Padre? > _"Wdrożenie AION znacząco uprościło i przyspieszyło nasz proces przygotowywania ofert eventowych w wielu aspektach. Doceniamy elastyczność rozwiązania i wsparcie zespołu wdrożeniowego. Każdego dnia odkrywamy kolejne zastosowania dla AION w naszej organizacji i czerpiemy z jego możliwości pełnymi garściami mimo iż wykorzystujemy zaledwie część jego potencjału."_ > > **Jakub Ćwikliński** > Wiceprezes zarządu w El Padre Ten ostatni punkt jest szczególnie interesujący - zespół El Padre wykorzystuje "zaledwie część potencjału" AION, co oznacza, że **korzyści będą rosły w czasie**, w miarę jak odkrywają nowe zastosowania platformy. ## Kluczowe wnioski: Co można wynieść z tego case study? ### 1. AI w branży eventowej to teraźniejszość, nie przyszłość El Padre pokazuje, że można zyskać **nawet 50% przyspieszenia procesów** - i to w branży, która wydawałaby się wysoce kreatywna i trudna do automatyzacji. ### 2. Integracja i centralizacja danych to klucz Bez zunifikowanej bazy wiedzy AI nie ma z czego czerpać. Pierwszy krok - uporządkowanie danych - był fundamentem sukcesu. ### 3. Wdrożenie nie musi być długie **6 tygodni wystarczyło**, by 25-30 osób pracowało efektywniej. Nie trzeba wielomiesięcznych projektów transformacyjnych. ### 4. ROI jest wymierny Oszczędności czasowe (75-120h miesięcznie) przekładają się bezpośrednio na możliwość obsługi większej liczby klientów bez zwiększania kosztów. ### 5. AI wspiera, nie zastępuje Zespół kreatywny nadal tworzy unikalne koncepcje - AI po prostu odciąża ich od żmudnych, powtarzalnych zadań. ## Dla kogo to rozwiązanie? Wdrożenie AI w procesach ofertowania sprawdzi się szczególnie w: - **Agencjach kreatywnych i eventowych** z dużą liczbą ofert - **Firmach, gdzie proces ofertowania jest złożony i czasochłonny** - **Organizacjach chcących skalować biznes** bez proporcjonalnego wzrostu zespołu - **Przedsiębiorstwach szukających przewagi konkurencyjnej** przez szybkość i innowacyjność Jeśli Twoja firma boryka się z podobnymi wyzwaniami - **przeciążonymi zespołami**, **czasochłonnymi procesami**, **brakiem centralizacji wiedzy** - warto rozważyć wdrożenie AI.

Chcesz zautomatyzować procesy ofertowania w firmie?

Pomogę Ci zbudować inteligentny system, który przyspieszy tworzenie ofert, wykorzysta wiedzę zespołu i ograniczy pracę manualną. Od analizy procesu przez projektowanie asystentów AI po wdrożenie i szkolenia.

Umów bezpłatną konsultację
## FAQ
### Jakie wyniki można osiągnąć dzięki automatyzacji procesu ofertowania z AI? Typowe rezultaty to 10-50% szybsze przygotowywanie ofert, 10-15% wzrost produktywności zespołu i oszczędność 75-120 godzin miesięcznie przy 15 ofertach. Pozwala to obsługiwać więcej klientów bez zwiększania zespołu i składać oferty na mniejsze projekty, które wcześniej były nierentowne.
### Ile trwa wdrożenie AI w procesie ofertowania? Typowe wdrożenie zajmuje około 6 tygodni w trzech etapach: integracja i centralizacja danych (tydzień 1-2), wdrożenie narzędzi AI (tydzień 3-4), implementacja workflow i szkolenie zespołu (tydzień 5-6). Nie trzeba wielomiesięcznych projektów transformacyjnych, by zobaczyć pierwsze efekty.
### Jak AI wspiera proces tworzenia ofert w agencjach kreatywnych i eventowych? Automatyczna transkrypcja spotkań z klientami i wyciąganie kluczowych informacji (budżet, wymagania, terminy), przeszukiwanie bazy archiwalnych ofert jako wzorców, automatyzacja kosztorysów i wycen, generator wstępnych wizualizacji. AI odciąża od żmudnych zadań, zespół kreatywny skupia się na unikalnych koncepcjach.
### Jak obliczyć ROI z automatyzacji procesu ofertowania? Zmierz średni czas przygotowania oferty przed i po wdrożeniu, pomnóż oszczędność przez liczbę ofert miesięcznie i stawkę godzinową zespołu. Przykład: 5-8h oszczędności × 15 ofert = 75-120h/miesiąc. Te godziny można przeznaczyć na więcej ofert lub projekty kreatywne bez dodatkowych kosztów.
### Dla jakich firm sprawdzi się automatyzacja ofertowania z AI? Agencje kreatywne i eventowe z dużą liczbą ofert, firmy ze złożonym i czasochłonnym procesem ofertowania, organizacje chcące skalować bez proporcjonalnego wzrostu zespołu. Kluczowe sygnały kwalifikujące: przeciążone zespoły, rozproszona wiedza w folderach, nierentowne mniejsze projekty.
--- # Automatyzacja poczty email z AI - Jak Frontdesk AI rewolucjonizuje obsługę klienta Source: https://pawel.lipowczan.pl/blog/automatyzacja-email-frontdesk-ai Published: 2025-11-10 # Automatyzacja poczty email z AI - Jak Frontdesk AI rewolucjonizuje obsługę klienta Obsługa zbiorczych skrzynek email to wyzwanie dla wielu firm. Setki wiadomości dziennie, powtarzające się pytania, konieczność szybkiej reakcji. Frontdesk AI rozwiązuje ten problem poprzez inteligentną automatyzację. ## Czym jest Frontdesk AI? Frontdesk AI to system do automatycznego przetwarzania i kategoryzacji poczty przychodzącej. Wykorzystuje sztuczną inteligencję (OpenAI) do analizy treści wiadomości, klasyfikacji według kategorii i automatycznego udzielania odpowiedzi. ## Główne funkcjonalności ### 1. Automatyczna kategoryzacja System analizuje treść każdej wiadomości i automatycznie przypisuje ją do odpowiedniej kategorii: - Zapytania o ofertę - Reklamacje - Pytania techniczne - Faktury i płatności - Spam ### 2. Inteligentne odpowiedzi ```text Dla najczęstszych pytań system automatycznie generuje odpowiedzi bazując na przygotowanej bazie wiedzy FAQ. ``` Każda odpowiedź jest kontekstowa i dopasowana do konkretnego pytania klienta. ### 3. Routing do odpowiednich osób Wiadomości wymagające ludzkiej interwencji są automatycznie przekierowywane do właściwych działów lub osób. ## Stack technologiczny - **Make** - automatyzacja workflow i integracje - **OpenAI GPT-4** - analiza treści i generowanie odpowiedzi - **Gmail/Outlook API** - integracja z pocztą email - **Airtable** - baza wiedzy i tracking wiadomości ## Przykładowy przepływ pracy 1. Nowa wiadomość email wpływa na skrzynkę 2. Make webhook przechwytuje wiadomość 3. OpenAI analizuje treść i intencję wiadomości 4. System sprawdza bazę wiedzy w Airtable 5. Jeśli znajdzie odpowiedź - automatycznie wysyła reply 6. Jeśli nie - przekierowuje do odpowiedniej osoby z kontekstem ## ROI i oszczędności Typowy klient oszczędza: - **20-30 godzin pracy miesięcznie** na obsłudze email - **90% redukcja czasu reakcji** na standardowe pytania - **100% dostępność** - system działa 24/7 - **Konsystentna jakość** odpowiedzi ## Wdrożenie Proces wdrożenia Frontdesk AI: 1. **Analiza** - mapowanie typowych kategorii wiadomości 2. **Baza wiedzy** - przygotowanie FAQ i szablonów odpowiedzi 3. **Konfiguracja** - ustawienie przepływów w Make 4. **Testy** - weryfikacja działania na próbnej grupie 5. **Go-live** - uruchomienie dla całej poczty 6. **Optymalizacja** - dostrajanie na podstawie feedbacku Typowy czas wdrożenia: 1-2 tygodnie. ## Podsumowanie Frontdesk AI to rozwiązanie dla firm, które: - Otrzymują wiele powtarzających się pytań - Potrzebują szybkiej reakcji na wiadomości - Chcą odciążyć zespół od rutynowych zadań - Dbają o jakość obsługi klienta

Chcesz zautomatyzować obsługę emaili w swojej firmie?

Pomogę Ci wdrożyć inteligentną automatyzację poczty email, która odciąży zespół od rutynowych pytań i przyspieszy reakcję na zapytania klientów. Od analizy przypadków przez konfigurację po testy i uruchomienie.

Umów bezpłatną konsultację
## FAQ
### Czym jest Frontdesk AI i jak automatyzuje obsługę poczty email? Frontdesk AI to system automatycznego przetwarzania poczty przychodzącej wykorzystujący OpenAI do analizy treści. System kategoryzuje wiadomości (zapytania, reklamacje, faktury, spam), automatycznie odpowiada na standardowe pytania z bazy wiedzy FAQ i przekierowuje złożone sprawy do odpowiednich osób z kontekstem.
### Jak działa automatyczna kategoryzacja wiadomości email przez AI? OpenAI GPT-4 analizuje treść i intencję każdej wiadomości, przypisując ją do zdefiniowanej kategorii. System sprawdza bazę wiedzy w Airtable - jeśli znajdzie odpowiedź, automatycznie wysyła reply. Jeśli nie, przekazuje wiadomość do właściwej osoby wraz z kontekstem i sugerowaną kategorią.
### Jakie oszczędności daje automatyzacja poczty email z AI? Typowy klient oszczędza 20-30 godzin pracy miesięcznie, redukuje czas reakcji na standardowe pytania o 90% i zyskuje całodobową dostępność (24/7). Dodatkowa korzyść: konsystentna jakość odpowiedzi - każdy klient otrzymuje ten sam standard obsługi niezależnie od pory dnia.
### Dla jakich firm sprawdzi się system automatyzacji emaili typu Frontdesk AI? Dla firm otrzymujących wiele powtarzających się pytań, potrzebujących szybkiej reakcji na wiadomości i chcących odciążyć zespół od rutynowych zadań. Idealne dla działów obsługi klienta, skrzynek info@ i support@, gdzie duża część zapytań dotyczy FAQ, statusu zamówień lub standardowych procedur.
### Ile trwa wdrożenie systemu automatyzacji poczty email? Typowy czas wdrożenia to 1-2 tygodnie. Proces obejmuje: analizę typowych kategorii wiadomości, przygotowanie bazy wiedzy i szablonów odpowiedzi, konfigurację przepływów w Make, testy na próbnej grupie, uruchomienie dla całej poczty i optymalizację na podstawie feedbacku.
--- # No-Code Lead Generation - Jak zbudować system generowania leadów bez programowania Source: https://pawel.lipowczan.pl/blog/no-code-lead-generation Published: 2025-11-05 # No-Code Lead Generation - Jak zbudować system generowania leadów bez programowania Pozyskiwanie leadów to jeden z najważniejszych procesów w każdej firmie B2B. Ręczne wyszukiwanie kontaktów to jednak czasochłonny proces. Pokażę Ci jak zbudować system, który robi to automatycznie. ## Problem: Ręczne generowanie leadów Typowy proces ręcznego pozyskiwania leadów: 1. Wyszukiwanie firm w Google (15-30 min na listę) 2. Odwiedzanie stron firm, szukanie kontaktów (5-10 min na firmę) 3. Weryfikacja adresów email (2-3 min na kontakt) 4. Wprowadzanie do CRM (1-2 min na rekord) **Efekt:** 2-3 godziny pracy = 10-15 zweryfikowanych leadów ## Rozwiązanie: Automatyczny Lead Generator System wykorzystujący: - **n8n** - workflow automation - **Snov.io** - email finder & verifier - **Apollo** - baza firm i kontaktów - **The Company API** - dane firmowe - **Airtable** - centralna baza leadów ## Jak to działa? ### 1. Definiowanie kryteriów wyszukiwania W Airtable tworzymy tabelę "Campaigns" gdzie definiujemy: - Branża (np. "e-commerce", "SaaS") - Wielkość firmy (pracownicy, przychody) - Lokalizacja - Stanowiska decydentów (CEO, CTO, Marketing Director) ### 2. Automatyczne wyszukiwanie firm ```text n8n workflow: 1. Trigger: Nowa kampania w Airtable 2. Apollo Search: znajdź firmy według kryteriów 3. The Company API: wzbogać dane o firmie 4. Filtrowanie: usuń duplikaty i nieaktualne dane 5. Zapis do Airtable ``` ### 3. Pozyskiwanie kontaktów decydentów Dla każdej firmy system: - Wyszukuje profile decydentów na LinkedIn (Apollo) - Znajduje adresy email (Snov.io) - Weryfikuje poprawność email (Snov.io) - Dodaje kontakty do bazy ### 4. Wzbogacanie danych System automatycznie dodaje: - Wielkość firmy - Przychody (jeśli dostępne) - Technologie używane na stronie - Aktywność w social media - Ostatnie newsy o firmie ## Architektura systemu **Baza Airtable:** - Tabela "Campaigns" - definicje kampanii - Tabela "Companies" - znalezione firmy - Tabela "Contacts" - decydenci w firmach - Tabela "Enrichment" - dodatkowe dane **n8n Workflows:** - Company Search (uruchamiany raz dziennie) - Contact Finder (ciągły, dla nowych firm) - Email Verifier (weryfikacja co 7 dni) - Data Enrichment (wzbogacanie danych) ## Efektywność **Ręcznie (8h pracy):** - 30-40 firm - 60-80 zweryfikowanych kontaktów - Koszt: 8h × stawka godzinowa **Automatycznie (24h):** - 500-1000 firm - 1500-3000 zweryfikowanych kontaktów - Koszt: ~$50 (API calls + narzędzia) **ROI: 95% redukcja kosztów pozyskania leada** ## Koszty operacyjne Miesięczne koszty przykładowego setup: - n8n (self-hosted): $0 - Snov.io (1000 credits): $39/mo - Apollo (Basic): $49/mo - The Company API (1000 calls): $29/mo - Airtable (Pro): $20/mo **Razem: ~$140/mo** vs kilkaset godzin pracy ręcznej ## Case Study: Agencja marketingowa **Przed:** - 2 osoby full-time na prospecting - ~200 leadów/miesiąc - Koszt: 2 × $3000 = $6000/mo **Po wdrożeniu:** - System automatyczny - ~2500 leadów/miesiąc - 1 osoba part-time na weryfikację - Koszt: $140 (tools) + $1000 (część etatu) = $1140/mo **Oszczędność: $4860/mo (81%)** ## Wdrożenie krok po kroku 1. **Tydzień 1:** Konfiguracja narzędzi i integracji 2. **Tydzień 2:** Budowa workflow w n8n 3. **Tydzień 3:** Testy i optymalizacja 4. **Tydzień 4:** Szkolenia i launch Czas wdrożenia: 3-4 tygodnie ## Podsumowanie Lead Generator to system który: - Działa 24/7 bez przerw - Generuje 10-15x więcej leadów - Kosztuje 80-90% mniej niż ręczna praca - Dostarcza wyższą jakość danych

Chcesz zautomatyzować generowanie leadów?

Pomogę Ci zbudować system automatycznej generacji leadów, który będzie działał 24/7, znajdował potencjalnych klientów i kwalifikował ich przed pierwszym kontaktem. Od koncepcji przez konfigurację po optymalizację.

Umów bezpłatną konsultację
## FAQ
### Z jakich narzędzi składa się system automatycznego generowania leadów no-code? Stack technologiczny obejmuje: n8n (workflow automation, self-hosted za $0), Snov.io (email finder i weryfikacja), Apollo (baza firm i kontaktów), The Company API (dane firmowe) oraz Airtable (centralna baza leadów). Wszystkie narzędzia integrują się bez pisania kodu przez API i gotowe konektory.
### Jakie oszczędności daje automatyzacja generowania leadów w porównaniu z pracą ręczną? 95% redukcja kosztów pozyskania leada i 10-15x więcej leadów. Ręcznie: 8h pracy = 30-40 firm i 60-80 kontaktów. Automatycznie: 24h działania systemu = 500-1000 firm i 1500-3000 zweryfikowanych kontaktów za ~$50 kosztów API. System działa 24/7 bez przerw.
### Jak działa automatyczne wyszukiwanie firm i kontaktów decydentów? Definiujesz kryteria w Airtable (branża, wielkość firmy, lokalizacja, stanowiska), n8n uruchamia workflow: Apollo wyszukuje firmy, The Company API wzbogaca dane, system filtruje duplikaty, Snov.io znajduje i weryfikuje emaile decydentów. Kontakty trafiają do centralnej bazy gotowe do outreach.
### Ile kosztuje miesięcznie automatyczny system lead generation? Około $140/miesiąc: n8n self-hosted $0, Snov.io (1000 credits) $39, Apollo Basic $49, The Company API (1000 calls) $29, Airtable Pro $20. Dla porównania: 2 osoby full-time na ręczny prospecting to koszt $6000/miesiąc przy znacznie mniejszej skali.
### Ile trwa wdrożenie systemu automatycznego generowania leadów? 3-4 tygodnie: konfiguracja narzędzi i integracji (tydzień 1), budowa workflow w n8n (tydzień 2), testy i optymalizacja (tydzień 3), szkolenia i uruchomienie produkcyjne (tydzień 4). Po wdrożeniu system wymaga minimalnej obsługi - część etatu na weryfikację jakości leadów.
--- # Chatboty oparte na AI - Od koncepcji do wdrożenia Source: https://pawel.lipowczan.pl/blog/chatboty-ai-od-koncepcji-do-wdrozenia Published: 2025-11-01 # Chatboty oparte na AI - Od koncepcji do wdrożenia Chatboty nowej generacji wykorzystujące LLM (Large Language Models) potrafią prowadzić naturalne, kontekstowe rozmowy. Nie są już ograniczone do sztywnych scenariuszy - rozumieją intencje i adaptują się do kontekstu. ## Czym różnią się od tradycyjnych chatbotów? **Tradycyjne chatboty:** - Sztywne scenariusze (decision trees) - Rozpoznawanie słów kluczowych - Brak rozumienia kontekstu - Ograniczona elastyczność **Chatboty AI:** - Naturalne zrozumienie języka - Kontekstowe odpowiedzi - Pamięć konwersacji - Wykonywanie akcji (booking, search, etc.) ## Architektura Context-based Chatbota ### 1. Frontend - Interface użytkownika **Web widget:** - Osadzany na stronie - Responsywny design - Opcje multimedialne (tekst, obrazy, przyciski) **Voicebot:** - Telefon (VAPI) - Voice interface na stronie - IVR integration ### 2. Backend - Logika konwersacji (n8n) ```text n8n Workflow: 1. Webhook receive message 2. Load conversation context 3. Search knowledge base (RAG) 4. Call LLM (OpenAI/Claude) 5. Execute actions if needed 6. Store conversation history 7. Return response ``` ### 3. Knowledge Base - Źródło wiedzy **RAG (Retrieval Augmented Generation):** Zamiast trenować model na swoich danych, używamy RAG: 1. Dokumenty są podzielone na chunks 2. Chunks są embedowane (wektory) 3. Zapisywane w vector database (Qdrant) 4. Przy zapytaniu: semantic search → top N chunks → context dla LLM **Zalety RAG:** - Aktualna wiedza (update bez retreningu) - Mniejsze koszty - Lepsze źródła odpowiedzi - Kontrola nad danymi ### 4. LLM - Mózg systemu **OpenAI GPT-4:** - Najlepsza jakość odpowiedzi - Function calling (akcje) - Koszt: ~$0.01 per 1k tokens **Claude 3.5 Sonnet:** - Świetne w analizie - Duży kontekst (200k tokens) - Koszt: ~$0.003 per 1k tokens ## Implementacja krok po kroku ### Krok 1: Przygotowanie bazy wiedzy Zbieramy dokumenty: - FAQ - Dokumentacja produktu - Artykuły blog - Polityki firmy Przetwarzanie: ```python # Podział na chunks (500-1000 tokenów) # Embedding przez OpenAI ada-002 # Zapis do Qdrant ``` ### Krok 2: Konfiguracja n8n workflow **Main conversation flow:** 1. Webhook trigger (user message) 2. Vector search w Qdrant (top 3 relevant chunks) 3. Format prompt z context 4. Call OpenAI z function calling 5. If function → execute & respond 6. Save to conversation history ### Krok 3: Function Calling - Akcje Chatbot może wykonywać akcje: ```json { "name": "book_meeting", "description": "Books a meeting with sales team", "parameters": { "date": "2025-11-20", "time": "14:00", "email": "user@example.com" } } ``` n8n wykrywa function call → integracja z calendly/Google Calendar → confirmation ### Krok 4: Testing i Optymalizacja - Test różnych promptów - Analiza failed conversations - A/B testing responses - Monitoring accuracy ## Case Study: automation.house **Wyzwanie:** Strona automation.house - dużo ofert (Note Taker, Lead Generator, etc.) Użytkownicy mieli trudności z wyborem odpowiedniego rozwiązania **Rozwiązanie:** Context-based chatbot który: - Zadaje pytania o potrzeby klienta - Rozumie kontekst biznesowy - Rekomenduje odpowiednie rozwiązania - Umawia konsultacje **Stack:** - n8n (hosting + workflow) - OpenAI GPT-4o (konwersacja) - Qdrant (baza wiedzy o produktach) - Airtable (tracking rozmów) **Wyniki:** - 40% wzrost engagement - 25% więcej umówionych konsultacji - 80% użytkowników kończy rozmowę z konkretną akcją ## Voicebots z VAPI VAPI to platforma do tworzenia voice AI: **Funkcje:** - Real-time voice conversations - Integration z telefonią - Transfer do człowieka - Recording & transcription **Use cases:** - Infolinia automatyczna - Kwalifikacja leadów przez telefon - Customer support 24/7 - Appointment booking ## Koszty wdrożenia **Setup (jednorazowo):** - Przygotowanie bazy wiedzy: 1-2 tygodnie - Konfiguracja workflow: 1 tydzień - Testing: 1 tydzień - **Razem: 3-4 tygodnie** **Miesięczne koszty operacyjne:** - n8n (self-hosted): $0-20 - OpenAI API (1000 rozmów): $30-50 - Qdrant Cloud: $25 - VAPI (voicebot): $99 - **Razem: $150-200/mo** vs. 1 pracownik customer support: $2500-3500/mo ## Best Practices 1. **Jasny cel konwersacji** - bot musi wiedzieć co ma osiągnąć 2. **Graceful degradation** - transfer do człowieka gdy nie wie 3. **Krótkie odpowiedzi** - nie pisz esejów 4. **Personality** - daj botowi charakter zgodny z brandem 5. **Testing** - testuj z prawdziwymi użytkownikami ## Podsumowanie Context-based chatboty to przyszłość customer experience: - Dostępność 24/7 - Konsystentna jakość - Skalowalność - Niski koszt operacyjny

Chcesz wdrożyć chatbota AI w swojej firmie?

Pomogę Ci zaprojektować, zbudować i wdrożyć chatbota dostosowanego do Twoich potrzeb biznesowych. Od analizy przypadków użycia przez konfigurację bazy wiedzy po integrację i optymalizację.

Umów bezpłatną konsultację
## FAQ
### Co to jest RAG i dlaczego jest lepszy od fine-tuningu modelu AI? RAG (Retrieval Augmented Generation) to technika, w której chatbot przeszukuje bazę wiedzy i podaje znalezione informacje jako kontekst dla LLM. Zalety nad fine-tuningiem: aktualizacja wiedzy bez kosztownego retreningu, niższe koszty, lepsza kontrola nad źródłami odpowiedzi i możliwość wskazania skąd pochodzi informacja.
### Ile kosztuje wdrożenie i utrzymanie chatbota AI dla firmy? Setup zajmuje 3-4 tygodnie (baza wiedzy, workflow, testy). Miesięczne koszty operacyjne dla 1000 rozmów: n8n self-hosted $0-20, OpenAI API $30-50, Qdrant Cloud $25, opcjonalnie VAPI dla voicebota $99. Łącznie $150-200/miesiąc vs $2500-3500 za pracownika customer support.
### Czym chatboty AI różnią się od tradycyjnych chatbotów opartych na słowach kluczowych? Tradycyjne chatboty działają na sztywnych scenariuszach (decision trees) i rozpoznają słowa kluczowe. Chatboty AI rozumieją naturalny język, pamiętają kontekst rozmowy, adaptują się do intencji użytkownika i wykonują akcje (rezerwacje, wyszukiwanie). Różnica to skala elastyczności - AI obsługuje zapytania, których twórca nie przewidział.
### Jak przygotować bazę wiedzy dla chatbota opartego na RAG? Zbierz dokumenty (FAQ, dokumentacja produktu, artykuły, polityki firmy), podziel je na chunks 500-1000 tokenów, wygeneruj embeddingi przez OpenAI ada-002 i zapisz w bazie wektorowej (np. Qdrant). Przy każdym zapytaniu chatbot wyszukuje 3-5 najbardziej relevantnych fragmentów jako kontekst dla odpowiedzi.
### Kiedy chatbot AI powinien przekazać rozmowę człowiekowi? Gdy nie zna odpowiedzi, użytkownik jest sfrustrowany, sprawa wymaga decyzji wykraczających poza uprawnienia bota lub dotyczy wrażliwych tematów (reklamacje, sprawy prawne). Graceful degradation to kluczowa best practice - bot informuje, że przekazuje do konsultanta, zamiast generować niepewne odpowiedzi.
--- # 5 Weak Spots of Claude Code and How to Fix Them Source: https://pawel.lipowczan.pl/en/blog/claude-code-weak-spots Published: 2026-07-13 Claude Code reads code, edits files and runs commands in the terminal. At those tasks it is very good. But ask it to watch a video or design a nice front end, and the limits show up fast. I test dozens of tools around Claude Code. The same set of five weak spots keeps coming back: video, design, memory, research and tokens (**token** - the unit of text a model counts and that you pay for). These are not flaws that break your work with code. They are areas where Claude Code simply has no built-in feature. The good news: each of these gaps can be patched. Not with theory, but with a concrete tool. For every area I will show what I use day to day. And I will be honest - some tools I have tested in daily work, some are only on my radar, and I will flag that clearly. The framing was inspired by a Chase AI video about five repos for Claude Code. But this is my own stack - tools I run myself, with my own caveats. ## 🎬 Video - Claude Code cannot see or make videos Claude Code cannot watch a video or generate one. By default it gets stuck on the transcript (**transcript** - a text record of the audio track). And a transcript is not enough when what matters is happening on screen. ### Watching: Claude Video `Claude Video` is a skill (**skill** - a ready set of instructions that extends the agent with a new ability) with the `/watch` command. I paste a URL or file plus a question. The agent pulls captions, extracts frames (**frame** - a single image from a video) and reads them as images. Before it answers, it has actually seen the video instead of guessing from the title. I use it for two things. I analyze other people's content faster than playback at 2x, and I pull knowledge that then lands in my notes base. It also works as a transcript grabber for YouTube. ```bash /watch https://youtu.be/dQw4w9WgXcQ co dzieje sie w 30 sekundzie? ``` The token cost depends mostly on frames, because each one is an image. That is why the skill has four modes, from cheapest to most expensive: ```text transcript - captions only, no frames (cheapest) efficient - keyframes, up to 50 balanced - on scene changes, up to 100 (default) token-burner - no frame cap (full coverage, pricey) ``` The `efficient` mode takes keyframes (**keyframe** - a frame the video itself marks as a reference point). When a video has no captions, the skill turns speech into text with the whisper model (**whisper** - a model that converts speech to text), for free through Groq. ### Creating: HyperFrames The opposite direction is generating video. `HyperFrames` builds video from HTML files - a composition is plain HTML with `data-*` attributes, no React and no custom language. The "brand" lives in my own CSS styles, so every video comes out consistent with my visual identity rather than generic. ```bash npx skills add heygen-com/hyperframes ``` It needs Node.js >= 22 and FFmpeg. I already wrote about generating video from code in the piece on [Remotion for explainer videos](/en/blog/remotion-explainer-videos-ai). The difference is simple: HyperFrames is HTML-native and under an Apache 2.0 license (commercial use with no limits), while Remotion ships distributed rendering in the cloud. ## 🎨 Design - no more generic AI slop Ask Claude Code for a page and you get exactly what everyone else gets: a hero section on top, rounded cards, a purple gradient. That is **AI slop** (**AI slop** - the generic look that instantly gives away a front end generated by a model). The reason is simple: models were trained on the same templates. I do not patch this with one tool, but with three skills lined up in order: who, system, guard. ### 1. UX RULER - who and why `UX RULER` forces the question "who is this for and what measurable value does it give", before I jump into features. It writes the decisions about audience and need into the repository as product memory (**product memory** - files in the repo that hold product decisions for the next human or agent). That way the choices are explicit and reusable. ### 2. UI UX Pro Max - a ready design system `UI UX Pro Max` generates a coherent design system (**design system** - a set of rules for color, typography and components) matched to the project type. A portfolio gets a different logic than a SaaS or a shop. I run it with Tailwind CSS and React as a base that I then refine, which saves prototype time. I covered this tool in more depth in [5 GitHub repos for Claude Code](/en/blog/5-github-repos-claude-code), so I will not repeat it here. ### 3. Impeccable - the design-language guard `Impeccable` is the layer that removes the tell-tale signs of AI: Inter everywhere, purple gradients, cards inside cards, gray text on colored backgrounds. It has a deterministic linter (**linter** - a tool that checks code or a layout against fixed rules) that runs with no model and no API key: ```bash npx impeccable detect src/ ``` The skill ships **23 commands** under `/impeccable`, among them `craft`, `shape`, `critique` and `colorize`, plus **27 deterministic rules** and an extra model review. This is the thing that keeps a front end from screaming "an AI made me". Together it forms a line: RULER says who for, Pro Max builds the system, Impeccable guards the language and cuts the AI slop. ## 🧠 Memory - Claude Code forgets after every session The end of a session is a reset. Claude Code has no built-in memory that accumulates across conversations. Every time it starts from zero. ### A second brain as memory My approach is a **second brain** (**second brain** - a text notes base that the agent reads, updates and searches). The repository is my memory: the agent writes knowledge into it, links notes and answers questions from the whole. This is a pattern that Andrej Karpathy named **LLM Wiki** (**LLM Wiki** - a base that the model itself builds and maintains from your sources). It differs from RAG (**RAG** - a technique where the model searches raw documents before answering). In RAG nothing accumulates, because the model rediscovers everything each time. In an LLM Wiki the knowledge stays in the files and grows. The key element is **progressive disclosure** (**progressive disclosure** - arranging notes so the agent can find them without cluttering the context window). The context window (**context window** - how much text the model sees at once) is limited, so a good index matters more than dumping everything into one file. I built this on Obsidian and Claude Code. The full architecture, 185 notes and three indexes I described in the [article on Karpathy's LLM Wiki](/en/blog/karpathy-llm-wiki-knowledge-base) and in [building a second brain in Obsidian](/en/blog/second-brain-obsidian-claude-code-skills). Where the code graph ends and the notes base begins, I lay out in the [piece on RAG](/en/blog/not-all-rag-is-equal). How to build such a base from scratch, I show step by step in the [free LLM Wiki course](/llm-wiki). ### NotebookLM-py - a fresh addition `NotebookLM-py` puts NotebookLM (**NotebookLM** - a Google engine where Gemini reads your sources and answers with citations) inside Claude Code. It is an unofficial API, so I use it at my own risk. ```bash uv tool install "notebooklm-py[browser]" notebooklm login ``` For now it mostly serves as a second route to YouTube transcripts. I am only starting, so I treat it as an addition, not a foundation. ## 🔎 Research - the built-in search is shallow The built-in search in Claude Code works, but only on the surface. What is missing is the middle ground between "shallow" and "a hundred agents and 10 million tokens on deep research" (**deep research** - a deep, multi-source study of a topic). ### Firecrawl - reliable page fetching When the built-in scraper (**scraper** - a tool that pulls the content of a page) fails, I reach for `Firecrawl`. It turns any page into clean markdown or structured JSON, handling JavaScript and pagination. ```bash pip install firecrawl-py ``` It has five modes: Scrape (one page), Crawl (a whole site by links), Map (a list of addresses), Search (search with content) and Extract (data by schema). It works as an MCP (**MCP** - Model Context Protocol, a standard for connecting tools to an agent), so it plugs straight into Claude Code. ### Deep research - to be tested `NotebookLM-py` from the memory section also has a deep research mode on Google's side, so in theory cheaper. I have not tested it seriously yet, so I treat it as a candidate, not an endorsement. ### Research swarm - structure wired into the wiki For structured research I use a skill that fires up an agent swarm (**agent swarm** - a swarm of independent agents working in parallel). One agent per item, in batches. Each writes a validated JSON record, and the finished report lands in my notes base. That turns a one-off research run into a permanent entry in the wiki. ## 🪙 Tokens - verbose output burns budget and context Long-winded answers (**output** - what the agent prints in its response) cost twice. You pay for tokens and you clog the context window. The longer the output, the faster the agent loses room for what matters. ### Caveman - output compression `Caveman` is a skill that makes the agent talk like a caveman: no articles, no filler, no pleasantries, just the meat. It cuts an average of **~65% of output tokens** (range 22-87%), keeping 100% of the technical content. ```powershell # Windows (PowerShell) irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex ``` ```bash # macOS / Linux / WSL curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash ``` It has three levels: `/caveman lite`, `/caveman full` and `/caveman ultra`. And here is my honest caveat: in `full` mode the descriptions can become unreadable, especially in Polish. So I stick to `lite`, and for code, commits and security writing I switch to `normal mode`. An important detail from the docs: Caveman only cuts the output, not the thinking tokens (**reasoning** - tokens the model spends on internal reasoning). It shrinks the mouth, not the brain. The biggest win is readability and speed, and the savings are a bonus. ### On the radar - candidates, not endorsements The rest of the set sits on my to-check list. I list them, but I flag clearly that I have not tested them myself yet. - **Headroom** - compresses what the agent reads (input and context) by 60-95%. The input side, where Caveman handles the output. - **Ponytail** - pushes the agent toward minimal code, so fewer lines and fewer tokens. - **Graphify** - turns code into a knowledge graph so the agent does not read file by file. ## Key takeaways Five weak spots, five concrete moves: 1. **Video** - add `/watch` before you read a bare transcript again. For generation, check HyperFrames. 2. **Design** - line up three skills in order: who (UX RULER), system (UI UX Pro Max), guard (Impeccable). 3. **Memory** - a text wiki maintained by the agent beats RAG at small scale. If you build from scratch, start with the course. 4. **Research** - Firecrawl when the built-in scraper fails, and deep research only when you truly need it. 5. **Tokens** - Caveman cuts the output, but in Polish pick `lite`. Treat the rest as candidates to test. The best part is that every one of these tools is free and open source. The worst that can happen is you run one, dislike it and uninstall it. Pick the area that hurts you most and start there.

Want to get more out of Claude Code in your team?

I will help you pick tools and skills that fit your way of working, roll them out and set up a repeatable process with the agent. And if you prefer to start on your own, you can build the agent's memory step by step in the free LLM Wiki course.

Book a free consultation
## Useful Resources - [Claude Video](https://github.com/bradautomates/claude-video) - the `/watch` skill that gives the agent sight. - [HyperFrames](https://github.com/heygen-com/hyperframes) - generating video from HTML files. - [Impeccable](https://impeccable.style/) - a design language plus a linter that removes AI slop. - [NotebookLM-py](https://github.com/teng-lin/notebooklm-py) - NotebookLM in the Claude Code terminal. - [Firecrawl](https://www.firecrawl.dev/) - any page turned into clean markdown. - [Caveman](https://github.com/JuliusBrussee/caveman) - token compression in the agent's answers. - [Free LLM Wiki course](/llm-wiki) - how to build the agent's memory (a second brain) from scratch. ## FAQ
### In which areas is Claude Code weak out of the box? Claude Code is great with code, files and commands, but it has five weak spots: it cannot see or make video, it generates generic design, it forgets after a session, it has shallow search and verbose output. Each of these gaps is patched with a separate tool or skill. These are gaps in features, not flaws in working with code.
### Can Claude Code watch a YouTube video? Not by default, but the Claude Video skill with the `/watch` command changes that. It takes a URL and a question, pulls captions and video frames and reads them as images, so it really "sees" the recording. When captions are missing, it turns speech into text with the whisper model. It also works well for quickly grabbing transcripts.
### How do you give an AI agent memory between sessions? The simplest way is a second brain, a text notes base in a repository that the agent reads, updates and searches. This is the LLM Wiki pattern: knowledge accumulates in files, unlike RAG where the model rediscovers everything from scratch. How to build such a base from scratch, I show step by step in the [free LLM Wiki course](/llm-wiki).
### What is "AI slop" and how do you avoid generic design? AI slop is the generic look of a page that instantly gives away a front end generated by a model: the same hero layout, rounded cards and purple gradients. You avoid it by giving the agent a real design system and a linter that catches these patterns. I use the UI UX Pro Max and Impeccable skills for this, the latter has the `npx impeccable detect` command.
### Does Caveman really save tokens and is it worth using? Yes, Caveman cuts an average of about 65% of output tokens while keeping the full technical content. The caveat: in `full` mode the descriptions can be unreadable, especially in Polish, so it is better to set `lite`. Remember that it only shortens the output, not the model's thinking tokens.
### Are these tools for Claude Code free? Yes, all five solutions are free and open source. Some need your own keys or add-ons: Claude Video may need a whisper key for videos without captions, and Firecrawl an account for the API mode. Installation usually comes down to a single command in the terminal.
--- # Not All RAG Is Equal: a Graph for Code, an Index for Notes Source: https://pawel.lipowczan.pl/en/blog/not-all-rag-is-equal Published: 2026-07-09 ## "RAG mostly means vector databases to me" - an objection that gets it right While running my [free LLM Wiki course](/llm-wiki) (course in Polish), I got a comment from a technical reader, which I paraphrase: "RAG mostly means vector databases to me. But for code there are more interesting approaches, based on syntactic analysis - codebase-memory-mcp, for example, which its authors call Hybrid LSP. Probably mediocre for notes, but for coding it works quite well." He's right. In both halves of that sentence. Instead of pushing back, I'd rather unpack it - inside that comment sits a map missing from most discussions of **RAG** (Retrieval-Augmented Generation - a technique where the model pulls content from an external source before answering, instead of relying only on what it memorized in training). "Add RAG to the project" sounds like a decision. It isn't one. It's like saying "use a database" without specifying relational, document or graph. By the end of this article you'll be matching the retrieval mechanism to the material: one for code, another for notes. You'll also know when plain grep (searching files for an exact string) beats everything else. ## RAG is a family of techniques, not one thing The "RAG = vector database" association comes from two years in which most deployments looked exactly like that. A vector database stores **embeddings** (numeric representations of text that let you search by similarity of meaning rather than exact words). That's one flavor of RAG, not its definition. Take a broader definition: RAG is any solution that cuts **token** usage (tokens are the units in which a model meters text) and answer drift by pulling in the right content. Four substrates fit: - **lexical** - grep and BM25 (a classic text-relevance scoring algorithm), searching by exact words; - **semantic** - vectors and similarity of meaning; - **structural** - graphs of symbols, definitions and calls, built from the code itself; - **curated index** - an agent-maintained table of contents for the knowledge base, as in [LLM Wiki](/en/blog/karpathy-llm-wiki-knowledge-base). My second brain (a knowledge base of markdown files, maintained by an agent) is RAG under this definition too. The mechanism just differs: instead of computing vector similarity, the agent reads a lightweight index and pulls 2-3 relevant notes. The argument is not about the goal. It is about the mechanism - and **the mechanism has to fit the material**. ## Code parses into symbols, notes don't Code differs from notes in one property that changes everything: its grammar is formal, and a parser resolves it unambiguously. The call `user.profile.name()` resolves exactly - you know where the definition lives, what type comes back, who else calls it. From material like that you can build an **AST** (abstract syntax tree - the structure a parser breaks code into), and from many trees a **call graph** (a map of "who calls whom" across the whole repository). That structure is precise and verifiable. Notes have grammar too - natural language grammar - but no machine resolves it unambiguously. The sentence "the decision was right because the client already ran PostgreSQL" does not decompose into symbols: no definitions, no types, no unambiguous references. That's why a curated index plus links between notes works for a knowledge base, and a syntax graph doesn't. Researchers from Duke and Snowflake put it most sharply in the HOMER paper on agent memory: **similarity ≠ causality**. Two fragments being close in embedding space is not the same as one depending on the other. Ask "who calls this function" and a vector answers with similarity-based guessing, while a graph answers with facts. My reader noticed the same thing from the practical side: "mediocre for notes, works for coding". Code has structure you can index without guessing - notes don't. ## Case study: codebase-memory-mcp The tool from the objection shows well what mature code retrieval looks like. `codebase-memory-mcp` (DeusData) is an **MCP** server (Model Context Protocol - the standard for plugging tools into an AI agent) that indexes an entire repository into a persistent knowledge graph of the code. Despite the "syntactic analysis" label, it's a hybrid of three layers: 1. **Symbol graph** - the tree-sitter parser (breaks code into ASTs, supports 158 languages) plus "Hybrid LSP", a lightweight reimplementation of what **LSP** (Language Server Protocol - the engine behind "go to definition" and "find references" in your editor) does: resolving types, imports and inheritance without running a full language server. 2. **Vectors** - built-in `nomic-embed-code` embeddings, no API key needed. 3. **Full text** - BM25 word search. The point: the tool uses vectors itself - it just doesn't start with them. It starts with structure and saves semantic similarity for fuzzy questions. The number in the README headline is striking: **3,400 tokens instead of 412,000** on the same task. Five graph queries instead of grepping the repository file by file - a **99.2%** reduction. One caveat: that's the authors' own claim (README plus an arXiv preprint), not an independent measurement. The direction matches what I know from my own note base, though: switching from "search everything" to "index first" cuts the cost of a single question from ~35k to ~3k tokens, roughly **30x**. Two different materials, two pairs of numbers, one pattern: structure before brute-force reading. ```text Question: "what breaks if I change the signature of update_profile()?" grep file by file: dozens of searches and reads -> ~412,000 tokens symbol graph: 5 structural queries -> ~3,400 tokens (numbers: claimed by the codebase-memory-mcp authors) ``` Honestly about the limits: full type resolution works for Python, TypeScript/JavaScript, Go, Rust, Java, C, C++, C#, Kotlin and PHP. For other languages the tool falls back to text matching - you'll get an answer, just a less precise one. ## A map of retrieval substrates codebase-memory-mcp is one point on a larger map. A survey of code-retrieval research (arXiv:2510.04905) splits the methods into **graph-based** (the core is an explicitly built graph of code structure) and **non-graph** (the repository treated as a bag of text, searched lexically or semantically). Six practical approaches line up along that axis: | Substrate | How it works | Examples | |-----------|--------------|----------| | AST graph (tree-sitter) | parse the code, build a graph of definitions and calls | codebase-memory-mcp, Graphify, Aider's repo map | | LSP as context | the agent asks a language server about symbols and types | Serena, mcp-language-server | | Persisted facts (SCIP/LSIF) | facts about code stored and queried like a database | Glean (Meta), srctx | | Agentic grep | the agent generates and runs its own searches, zero index | Claude Code, Cline, GrepRAG | | Vectors (semantic) | embedding similarity | classic RAG | | Curated index next to code | a tree of `AGENTS.md` files; the agent descends from root to the edit site | DOX | **SCIP and LSIF** are formats for persisting facts about code (who defines, who uses), storable in the repository and shareable across a team. The sixth row needs a word of explanation: `AGENTS.md` is an instruction file a coding agent reads automatically when it starts working in a repository. The DOX pattern (the `agent0ai/dox` repo) multiplies it - one `AGENTS.md` per folder, with the folder's purpose, local conventions and an index of subfolders - and the agent walks that tree from the root down to the edit site, updating the docs after each change. It's the same curated index as for notes, just versioned together with the code. There are more approaches - execution-path graphs, hybrids with team knowledge - but these six cover most decisions you'll actually make. ## Honestly: when grep beats the graph If structure always won, this article would be shorter. It doesn't win. The GrepRAG preprint (Zhejiang University, January 2026) flips the perspective. Instead of building any index, the model generates around 10 `ripgrep` commands itself, runs them against the raw repository and re-ranks the hits with BM25 weights. Zero index means zero stale index - nothing drifts out of sync with the code. According to the authors, the approach matched or beat graphs and vectors on code completion, with retrieval ~13x faster on average (35x in the extreme case). It's a preprint with self-reported results - I treat it as a claim, not a fact. The stronger argument came from practice. The Claude Code team dropped an early version built on RAG with a local vector database, because agentic search - grep plus file navigation - worked better and had no stale-index problem. Cline, another popular coding agent, published the same position in its "why we don't index your codebase" manifesto. On the other side stands RepoGraph (ICLR 2025, peer-reviewed): adding a repository graph improved four agents' results by an average of **32.8%** on SWE-bench-Lite (a benchmark of fixing real bugs from GitHub projects). Grep versus graph - who's right? Both camps. The task decides: | Task | Winner | Why | |------|--------|-----| | Local completion of a fragment | agentic grep | cheap, no index, no drift - the answer is nearby | | A change crossing many files | graph / LSP | you need the map of calls and types that a local window won't show | | "Who calls X? What breaks after a change?" | LSP / definition graph | an exhaustive result, not a ranked list of candidates | | Fuzzy "where's the login logic" | vectors | matching intent, not exact structure | So in practice you don't choose "graph or grep". Grep takes the local work, structure takes changes that cross module boundaries, vectors take fuzzy questions about intent. codebase-memory-mcp quietly concedes the same: one tool ships a graph, embeddings and graph-augmented grep. One gap I have to note. All these tools are verified on Python, TypeScript, Go or Rust. For C# and Roslyn (the .NET compiler) I found no confirmed deployment of any of them. The LSP bridges should work in theory; in practice that's a test to run, not a fact to cite. I barely write code by hand anymore - agents do it - but that doesn't void the problem: the retrieval an agent uses to read a repository still depends on language support. If your team lives in .NET, test before you trust the benchmarks. ## An agent's two layers of memory Back to the objection. "Structural for code, mediocre for notes" - agreed. But the conclusion isn't "pick one". In a coding agent you want both layers at once, because they remember different things. **The graph and LSP remember what the code looks like**: what calls what, where the definition lives, what breaks after a signature change. **The second brain remembers why the code looks that way**: which convention we adopted, what we tried before and what failed, why we picked this library and not that one. The graph will give you the full call tree. It won't tell you that we enforced the current convention after the previous one blew up production in March. That information lives in no AST, because nobody wrote it there. It stayed in people's heads, in Slack threads or - if you keep a knowledge base - in a single note with a date and a rationale. Code doesn't record intent. A brain does. Where does the `AGENTS.md` tree live in this split? In between. Local folder conventions - "tests here are written like this, don't touch that module" - belong right next to the code, versioned with it, and that's exactly what the DOX pattern does. The brain keeps, one floor higher, what no single repository can hold: decisions and standards shared across projects, the history of approaches, lessons from failures. Two layers become three: the graph remembers structure, `AGENTS.md` remembers local conventions, the brain remembers knowledge above the project. That's why my reader was right in both halves of his sentence - and his objection describes not competition but a division of labor between the memory layers of one agent. Keep the code's structure in a graph. Keep decisions, conventions and lessons in a note base - [portable and agent-readable](/en/blog/okf-standard-portable-knowledge-base). How to build one from scratch is what I show in my [free LLM Wiki course](/llm-wiki) (in Polish). ## Key Takeaways 1. **Not all RAG is equal.** Pick the substrate: lexical (grep/BM25), semantic (vectors), structural (AST/LSP/graph) or a curated index. A vector database is one flavor, not the definition. 2. **Code parses into symbols, notes don't.** Default to structure for changes crossing file boundaries; use an index and links for notes. 3. **Grep is a contender, not a backup.** No index beats a stale index for local work - Claude Code and Cline built their whole approach on that. 4. **Match the mechanism to the task, not the fashion.** Treat README and preprint benchmarks as the authors' claims until someone reproduces them independently. 5. **Give a coding agent both layers of memory.** The graph remembers structure, the brain remembers intent, and local folder conventions belong next to the code (an `AGENTS.md` tree). Only together do they answer both "what breaks" and "why it's like this".

Wondering what retrieval to feed your AI agent?

I help teams pick their agents' memory layers for real projects: a graph for code, a knowledge base for decisions and conventions. I'll tell you what will work for your team before you buy another tool.

Book a free consultation

You can build the knowledge layer yourself - my free LLM Wiki course (in Polish) shows how, step by step. The course template pairs with a code repository.

## Useful Resources - [codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp) - a code graph as an MCP server; preprint: [arXiv:2603.27277](https://arxiv.org/abs/2603.27277) - [Retrieval-Augmented Code Generation - a survey](https://arxiv.org/abs/2510.04905) - the graph vs non-graph taxonomy - [GrepRAG](https://arxiv.org/html/2601.23254v2) - index-free retrieval: the model generates its own ripgrep commands - [RepoGraph](https://arxiv.org/abs/2410.14684) - peer-reviewed evidence that a repository graph improves agent results - [Cline: why we don't index your codebase](https://cline.bot/blog/why-cline-doesnt-index-your-codebase-and-why-thats-a-good-thing) - the agentic-grep manifesto - [Serena](https://github.com/oraios/serena) - LSP as agent tools (symbols, references, refactoring) - [DOX](https://github.com/agent0ai/dox) - a self-documenting `AGENTS.md` tree: a curated index versioned with the code - [Aider's repo map](https://aider.chat/docs/repomap.html) - the ancestor of repository maps with a token budget - [How Karpathy's LLM Wiki helped me organize my knowledge base](/en/blog/karpathy-llm-wiki-knowledge-base) - a curated index in practice - [I built a second brain. It matched Google's standard (OKF)](/en/blog/okf-standard-portable-knowledge-base) - a portable note base ## FAQ
### Is RAG always a vector database and embeddings? No. A vector database is one of four retrieval substrates, next to lexical (grep, BM25), structural (AST/LSP graphs) and a curated index. If you define RAG as pulling the right content before the model answers, then LLM Wiki with its index is RAG too - just without vectors.
### When is grep enough for searching code, and when do you need a symbol graph? Grep is enough for local work: completing a fragment, a known string, an answer near the edit site. You need a graph or LSP for changes crossing file boundaries and for questions like "who calls X, what breaks after a change". There grep returns false positives and vectors guess.
### How does retrieval for code differ from retrieval for notes? Code has a formal grammar that a parser resolves unambiguously, so you can build an exact graph of definitions, types and calls from it. A note's text doesn't decompose into symbols - the sentence "why we made this decision" has no definitions or types. That's why a curated index and links between notes work for a knowledge base, and structure works for code.
### Does LLM Wiki replace tools like codebase-memory-mcp? No, they're complementary memory layers for an agent. The code graph remembers structure: calls, types, dependencies between files. LLM Wiki remembers intent: decisions, conventions, rejected approaches. A coding agent uses both at once.
### What does "similarity ≠ causality" mean for vector databases? Two fragments being close in embedding space doesn't mean one depends on the other. On structural questions like "who calls this function", a vector answers with similarity - guessing - while a call graph answers with facts. The phrase comes from the HOMER paper (Duke + Snowflake) on agent memory.
### Where should project conventions live - in AGENTS.md files or in a second brain? Keep local folder conventions (test style, modules that must not be touched) next to the code - in an `AGENTS.md` tree versioned with the files, as in the DOX pattern. Keep what crosses a single repository in the second brain: architectural decisions with rationale, standards shared across projects, the history of rejected approaches. The symbol graph complements both - it remembers the structure no markdown describes.
--- # I Built a Second Brain. It Already Matched Google's Standard (OKF) Source: https://pawel.lipowczan.pl/en/blog/okf-standard-portable-knowledge-base Published: 2026-06-21 ## I built a second brain. It turned out to be 100% on Google's standard (OKF) A while ago I wrote about [how Karpathy's LLM Wiki helped me organize my knowledge base](/en/blog/karpathy-llm-wiki-knowledge-base) - over 300 notes managed by an agent, three navigation indexes, workflows for ingest and compilation. The system works. Every day. It's my second brain. But recently I caught myself on an uncomfortable question. All that knowledge lives in **my** tools - in Obsidian, in Quartz, in my `CLAUDE.md`. So what if I want to hand it to someone else? Switch tools? Or treat it as a company asset meant to outlive any program and any model? Notes locked in someone's private format are vendor lock-in - except self-inflicted. The conclusion I reached is simple: **durability of knowledge isn't the tool - it's the format.** That's when I came across **OKF (Open Knowledge Format)** by Google. I audited my brain against this standard, expecting a sizable gap. The result surprised me: compliance around 100% - even though I never designed the brain for OKF. In this article I'll show why that happened, what exactly I checked, and what happens when plain markdown isn't enough. No theory - a concrete audit and concrete numbers. ## What OKF (Open Knowledge Format) is `OKF` is an open format for recording knowledge, designed for the agent era. Technically there's no magic in it: it's **markdown + YAML frontmatter**. The whole philosophy fits in one tagline of the pattern - *„authored by people, generated by agents"*. Exactly the split I'd arrived at organically. The base unit is a `Knowledge Bundle` - a directory of `.md` files. It can be a git repository, a tarball, or a subdirectory. And that's the heart of portability: `git clone` and you have the whole thing. Inside, a `Concept` is a single thought - **one `.md` file = one concept** (frontmatter plus markdown content). The format reserves two filenames: `index.md` (a directory listing for [progressive disclosure](/en/blog/karpathy-llm-wiki-knowledge-base)) and `log.md` (a chronology of changes, newest on top). The most interesting part is the minimal barrier to entry. **The only hard-required frontmatter field is `type`.** Everything else - `title`, `description`, `resource`, `tags`, `timestamp` - is recommended but optional. A producer may add its own keys, and a consumer **must** tolerate unknown fields and broken links. The simplest compliant concept looks like this: ```yaml --- type: note --- Note content in plain markdown. ``` It's worth understanding the difference between OKF and `LLM Wiki`. **OKF is an interoperability specification** - it says „how to record knowledge so it's exchangeable". **LLM Wiki is a working methodology** - it says „how an agent should build and maintain that knowledge". Two complementary layers, not competitors. One honest note. The OKF repository carries an explicit disclaimer: *„This repository and its contents are not an official Google product"*. OKF is an open interoperability spec, not a product commitment from Google. I'm writing this deliberately - credibility comes from precision, not from overstating. ## My brain matched the standard - even though I never designed it for it This is the heart of the story. I ran a real compliance audit of the brain against OKF v0.1. The verdict in one sentence: **the brain passes 100% of the hard conformance conditions of OKF v0.1; the discrepancies are purely in the recommended layer - cosmetics.** The spec defines three hard conformance conditions. The brain meets all three: | # | OKF requirement | Status | Evidence in brain | |---|-----------------|:------:|-------------------| | 1 | Every non-reserved `.md` has parsable YAML frontmatter | ✅ | Every note opens with `--- ... ---` | | 2 | Every frontmatter has a non-empty `type` | ✅ | `type: basic-note \| book-note \| knowledge-note \| tool \| compiled-note \| answer-note` | | 3 | Reserved files have the proper structure *when they exist* | ✅ | Brain doesn't use literal `index.md`/`log.md` → that role is filled by `_indexes/*`; the „when they exist" condition is met | The numbers out loud: **hard compliance = 100% (3/3)**. The recommended layer (all those optional fields and conventions) is **~85-90%**. And here it gets interesting, because all four discrepancies are naming-and-syntax issues - they're about *what I named something*, not *how I model knowledge*: | Area | OKF | brain | Nature | |------|-----|-------|--------| | Links | `[label](/path)` | `[[wikilinks]]` (Obsidian) | syntax; **Quartz compiles to `[label](path)` at build** | | Index | literal `index.md` | `_indexes/vault-map.md` + `catalog.md` + `graph.md` | name, not function | | Log | `log.md` | a `## Recent Changes` section in `_indexes/vault-map.md` + git history | role in a section, not a separate file | | Frontmatter keys | `description` / `resource` / `timestamp` | `summary` / `source` / `date` | 1:1 mappable on export | Look at the „Nature" column. Nowhere does it say „I'm missing this concept" - everywhere it's „I named it differently" or „the tool compiles this for me". Obsidian wikilinks are full-fledged OKF links after the Quartz build. `_indexes/` plays exactly the role of `index.md`. The fields map one to one. > „Compliance isn't a coincidence - the brain and OKF grow from the same LLM Wiki concept by Karpathy. OKF is an interoperability spec, the brain is a working implementation. The differences are tool choices, not differences in the knowledge model." This is **convergent evolution**. Two systems built independently - a spec inside Google and my Obsidian vault - converged on the same shape because they grew from the same idea. And that's precisely why the standard is credible: it wasn't invented from behind a desk, it was written down from where practice already heads. ## Why the format matters You might ask: if my system works, why do I need a standard at all? The answer is five concrete things the format gives you - and that no single tool can: - **Portability.** *„If you can `git clone` it, you can ship it."* A bundle is a repository - you copy the whole thing and it works for anyone, no install, no config, no dependence on my setup. - **Interoperability.** A knowledge bundle can be exchanged regardless of tool. Export from my brain, import into someone else's. This is a real future product direction - an extract of a knowledge base as a ready OKF bundle. - **Durability and future-proofing.** Tools die, the format stays. Markdown plus frontmatter will outlive Obsidian, outlive Quartz, outlive a specific LLM. In five years you'll still open these files. - **Dual readability.** The same file is read by an agent and a human. Without any tool you just open it in an editor - it's not a base in a closed binary format. - **An asset, not a silo.** Knowledge recorded in a standard becomes a resource you can audit, hand off, and value. And that's the bridge to the next question - what happens when this asset grows to company scale. ## When plain markdown isn't enough - Google Knowledge Catalog Here I switch to a cautious tone. I have **not tested** this part - I'm signaling a direction, not issuing a recommendation. Karpathy himself sketches the scale where this approach works: _„moderate scale (~100 sources, ~hundreds of pages)"_ - before embedding-based search becomes necessary. My brain is on the order of 300 notes today, still in that range: plain markdown with indexes handles it without embeddings and without RAG infrastructure. In the [previous article](/en/blog/karpathy-llm-wiki-knowledge-base) I showed this mechanism in numbers: progressive disclosure gives about **30x less context** per query than dumping the whole vault. This isn't a hard limit - more an order of magnitude up to which the index-first approach stays comfortable. For a personal second brain it's plenty. But above that - millions of documents, structured and unstructured data at once, many agents working in parallel - it's a different league. This is where **Google Knowledge Catalog** comes in (managed, per the product page „formerly Dataplex"). Google describes it as a _„universal context engine for your enterprise"_ - „always-on context and governance for your agents". Instead of your notes, it catalogs the whole data estate: it automatically harvests metadata from BigQuery, AlloyDB, Spanner, or Looker, and turns unstructured sources - PDFs, contracts, wikis - into a _„structured knowledge graph"_ that an agent can query. The same repository (`GoogleCloudPlatform/knowledge-catalog`) that contains the `okf/` directory also holds `samples/` and `toolbox/`. There's continuity here, not a leap. The same idea - knowledge standardized, semantic, legible to agents - just at enterprise level. OKF and LLM Wiki are personal scale; Knowledge Catalog is company scale. One chain, two ends. **Where exactly is the boundary?** It's not „local vs cloud" - my brain runs in the cloud too. The difference is qualitative, along four axes: - **What you catalog.** Brain = your authored notes. Catalog = a semantic layer over the company's entire, living data estate: tables, warehouses, files, BI models. - **Retrieval engine.** Brain = index-first plus progressive disclosure, no embeddings. Catalog = semantic search with sub-second latency over a knowledge graph of millions of entities. - **Governance.** Brain = git and conventions. Catalog = access policies (IAM), data quality, lineage, audit - retrieval respects permissions, so an agent only sees what it's authorized to. - **Freshness and concurrency.** Brain = a snapshot you maintain yourself. Catalog = always-on, updates with the data, and serves many agents at once through Context APIs and MCP tools. The file count is just a symptom. The real boundary is **data type plus retrieval engine plus governance plus dynamics**. For personal knowledge in prose, plain markdown wins on simplicity; for a company's heterogeneous data estate you need a layer like Knowledge Catalog. And once more, because it matters: the OKF repository bears the note *„not an official Google product"*, and I describe Knowledge Catalog here strictly as a direction I haven't deployed myself. No promises about product features. ## What follows from this 1. **Format beats tool.** If you're building a knowledge base, design it for a standard from day one. Migrating later costs - compliance from the start is free. 2. **Compliance can be convergent.** If your system grows from a good idea (LLM Wiki), you have a shot at hitting the standard without aiming for it. That's good news: you probably don't need to rewrite your brain, just export it. 3. **Knowledge in a standard is an asset.** Portable, auditable, exchangeable - and scaling from a personal brain up to Knowledge Catalog. If you want to start on a ready foundation, I share a **[`second-brain-template`](https://github.com/plipowczan/second-brain-template)** - „Use this template", `/onboard`, first ingest, and you have your own standard-compliant system in minutes. I'm also working on broader materials about building a second brain - if the topic interests you, it's worth keeping an eye out.

Want a knowledge base that's yours forever - portable and ready for agents?

I'll help you design a knowledge base architecture aligned with the standard (OKF) - from note structure and indexes to company scale. You can also grab the free template and spin up your own system in minutes.

Book a free consultation
## Useful Resources - [How Karpathy's LLM Wiki helped me organize my knowledge base](/en/blog/karpathy-llm-wiki-knowledge-base) - the base article on LLM Wiki, progressive disclosure, and the brain's architecture - [Second brain in Obsidian and Claude Code](/en/blog/second-brain-obsidian-claude-code-skills) - how to start from scratch - [brain.lipowczan.pl](https://brain.lipowczan.pl) - the live instance of my second brain - [second-brain-template](https://github.com/plipowczan/second-brain-template) - a ready template to start with - [OKF / knowledge-catalog (repo)](https://github.com/GoogleCloudPlatform/knowledge-catalog) - the OKF spec in the `okf/` directory (note: „not an official Google product") - [Google Knowledge Catalog](https://cloud.google.com/products/knowledge-catalog) - a managed data catalog platform (enterprise scale) ## FAQ
### What is the Open Knowledge Format (OKF) and who is behind it? OKF (Open Knowledge Format) is an open specification for recording knowledge based on markdown plus YAML frontmatter, designed for working with AI agents. The spec lives in the `GoogleCloudPlatform/knowledge-catalog` repository, in the `okf/` directory. Important note: the repository carries a „This repository and its contents are not an official Google product" disclaimer - it's an open interoperability standard, not a product commitment from Google.
### How does OKF differ from Karpathy's LLM Wiki? OKF is an interoperability specification - it defines how to record knowledge so it's portable and exchangeable between tools. LLM Wiki is a working methodology - it describes how an agent should build and maintain a knowledge base. They're complementary: OKF is about the format, LLM Wiki about the process. A working knowledge base uses both layers at once.
### The only required field in OKF is `type` - what does that mean in practice? It means a minimal barrier to entry: a compliant OKF file only needs a non-empty `type` field in its frontmatter, while everything else (`title`, `description`, `tags`, `timestamp`) is optional. A producer can add any custom keys, because a format consumer must tolerate unknown fields and broken links. This makes the standard liberal on write and strict only where it has to be.
### Is my existing Obsidian knowledge base compliant with OKF? Most likely in large part yes, especially if you use `.md` files with frontmatter. In my audit the brain passed 100% of the hard conformance conditions of OKF v0.1 (3/3), and the discrepancies were purely cosmetic. Obsidian `[[...]]` wikilinks compile through Quartz to `[label](path)` links, and fields like `summary` or `date` map one to one onto `description` and `timestamp` on export.
### When does plain markdown stop being enough and Google Knowledge Catalog steps in? The index-first approach on plain markdown works great at personal scale - on the order of hundreds of notes (my brain is a bit over 300 today), without embeddings and without RAG infrastructure. It's not a hard limit, just an order of magnitude. Above that - with millions of documents, structured and unstructured data, and many agents at once - you reach enterprise scale, which is what Google Knowledge Catalog (managed, „formerly Dataplex") addresses. I should stress, though, that I haven't tested this threshold personally - I'm signaling a direction, not a product recommendation.
--- # Software 3.0: why your app shouldn't exist Source: https://pawel.lipowczan.pl/en/blog/software-3-0-agentic-engineering-guide Published: 2026-05-29 The man who coined the term **vibe coding** said he'd "never felt more behind as a programmer." That's Andrej Karpathy - OpenAI co-founder, former head of AI at Tesla. In December, agentic models crossed a threshold for him: chunks of code "just came out fine," so he stopped correcting them and started trusting the system. This isn't a story about writing code faster. It's a story about programming turning into prompting - and if you read "software" as every digital business, it means your product is being rebuilt to be AI-first, whether you like it or not. For the last few months I've been building with agents every day - the entire portfolio you're reading this on, plus systems in two companies. The more I work with them, the more clearly I see this isn't "just another AI hype cycle." It's a change in the map: abstraction rises, and with it shifts what the human actually does. This article is an attempt to draw that map - so you know which paradigm you operate in, where your moat is, and what can't be delegated. ![The Karpathy Paradigm diagram - five bands: ① three paradigms Software 1.0→2.0→3.0 (rising abstraction), ② vibe coding raises the floor vs agentic engineering preserves the ceiling, ③ verifiability and jagged intelligence (refactor 100k-line vs car wash), ④ what humans still own - taste, judgment, understanding, ⑤ build for agents: sensors, actuators, data legible to LLMs](/images/karpathy-paradigm-software-3-0.webp) This diagram is the spine of the whole text - I return to each of the five bands in the sections that follow. ## Three paradigms: how you "program" the machine Karpathy arranges the evolution of software into three paradigms - and the key is *how* you pass your intent to the machine. - **Software 1.0** - you write explicit code, rule by rule. Step one, step two, step three. Brittle, not self-healing: the moment something falls outside the script, it breaks. - **Software 2.0** - you "program" through **data**. You don't write rules; you curate datasets and train neural-network weights. The logic lives in the weights, not in `if` statements. - **Software 3.0** - **prompting**. The context window is your lever over the interpreter that is the LLM. > "Software 3.0 is kind of about your programming now turns to prompting. And what's in the context window is your lever over the interpreter that is the LLM." - Andrej Karpathy And here comes the reframe that turns this from "a curiosity for programmers" into something that concerns everyone who builds a product. Read "software" as **every digital business**. The LLM is the engine - and you design the car around it. > "All businesses are literally being restructured to be AI first... AI in the middle being the engine and driving them forward. But that is all the engine is. You get to design your own car." This isn't abstract. You see the same thing in how companies are starting to operate as agentic systems - with context as the durable asset and software as a layer you regenerate whenever a better model lands. ## Four examples that make it concrete The paradigm sounds abstract until you see what it does to specific business models. Karpathy gives four examples, and each is a different type of product that's shifting right now. 1. **The installer.** Instead of a ballooning script ("step one: accept the terms, step two: find the folder..."), you summon an agent with one command. It *owns the goal* ("get this installed") and loops through problems it has never seen before. It's a small, self-healing "skill" - a blob of text you paste to your agent. 2. **The app (MenuGen).** Karpathy built a 2.0 app: you photograph a menu, it renders images of the dishes. It became "instantly useless" - because you can just hand the photo to Gemini and say "overlay the items using Nano Banana." Nothing to install. *"That app shouldn't exist."* And the same is coming for almost every single app. 3. **The course / specialized knowledge.** From sequential Udemy lectures, through a course reshaped by engagement data, to a coach-agent that *does it with you* while you take action. Not "ask me a question," but "build it with me." 4. **The service (video editing).** From manual Premiere Pro, through Descript trimming the silences, to a text box: "edit this in a viral MrBeast style, under 8 minutes" - and a finished render in minutes. The common denominator is brutally simple: **Software 3.0 sells the outcome, not the tool.** If your product is a tool that performs a step rather than delivering a result - it's in the crosshairs. ## Vibe coding raises the floor. Agentic engineering preserves the ceiling. This is, in my view, the most important new distinction in the whole interview - and the place most people get it wrong. **Vibe coding** and **agentic engineering** are two different things. **Vibe coding raises the FLOOR.** Anyone can now build software. It's democratizing and genuinely great - I described the whole approach in my [vibe coding guide](/en/blog/vibe-coding-guide). The floor rises for everyone. **Agentic engineering preserves the CEILING.** It's a real engineering discipline: keep the professional quality bar - no new vulnerabilities, it's *still your* software, you're responsible for it - **while** going faster. Agents are "spiky entities": fallible, a little stochastic, but extremely powerful. The craft is coordinating them without dropping the bar. > "Vibe coding is about raising the floor for everyone... agentic engineering is about preserving the quality bar of what existed before in professional software." And here's the other side of the coin. The ceiling is very high. The old "10× engineer" is now magnified *far* beyond 10× - for people who are good at this. It's not a zero-sum game between "AI" and "the human." It's a lever that widens the gap between someone who can direct agents and someone who just clicks "accept." ## Verifiability - the new constraint Where do these models' strengths and weaknesses come from? Karpathy gives one of the best explanations I've heard, and it comes down to a single word: **verifiability**. > "The previous generation of computers automated what you could specify. This generation automates what you can verify." Frontier models are trained in giant RL (reinforcement learning) environments with verification rewards. That's why their capabilities are **jagged**: peaks where the result can be verified (code, math), valleys elsewhere. Jaggedness = what's *verifiable*, **plus** what the labs happen to choose to train on. > "How is it possible that a state-of-the-art model will refactor a 100,000-line codebase or find zero-day vulnerabilities, and yet tell me to walk to a car wash 50 metres away?" For a founder, there's a concrete takeaway. A **verifiable** problem is tractable today - you can throw RL at it, build your own RL environments, or fine-tune. That's your moat, independent of what the big labs happen to prioritize. If you can cheaply and automatically check whether the output is good, you have leverage that competitors without that verification don't. ## What you can't outsource This is the emotional core of the whole thesis. If an agent can refactor 100,000 lines, what actually remains on the human side? What remains is **taste, judgment, and oversight**. You hold the spec, the plan, and the top-level design - the agent "fills in the blanks." What also remains is **fundamentals over API trivia**. Karpathy no longer remembers whether it's `keepdim` or `axis`, `reshape` or `permute` - the "intern" has perfect recall for that. But you still have to understand what's happening underneath (e.g. how tensor views and storage work) so you're not silently copying memory in the wrong place. And at the very bottom there's one thing you'll never hand off: > "You can outsource your thinking, but you can't outsource your understanding." Karpathy says outright that he now feels like the **bottleneck** - he's the one who knows what to build and why, and he's the one directing the agents. An anecdote about unreliability illustrates it well: the MenuGen agent tried to match users by cross-correlating **email addresses** from Stripe and Google instead of using a persistent user ID. It's a "why would you ever do that" mistake - the kind a human has to catch, because the agent lacks the judgment to notice it makes no sense. I see the same thing in my own work: the agent speeds everything up, but I'm the reviewer who decides what's correct. Building your own knowledge base that an agent maintains is exactly an investment in **understanding** - I described that process in my piece on a [Karpathy-style knowledge base](/en/blog/karpathy-llm-wiki-knowledge-base). ## Build for agents, not for humans If agents are becoming the primary "user" of your systems, the logical consequence is that you have to start designing *for them*. Karpathy has a pet peeve about this: documentation is still written for humans. "Why are people still telling me what to do? What is the thing I should copy-paste to my agent?" The direction is this: - Describe systems to **agents first**. - Decompose work into **sensors** (read the world) and **actuators** (act on the world). - Keep data structures **legible** - readable for the LLM. Ultimately we're heading toward a world where "my agent talks to your agent" - to arrange a meeting, say. And if so, even **hiring** needs to be refactored. Instead of algorithmic puzzles, you give a candidate a big project - e.g. "build a secure Twitter-clone for agents, then have 10 agents try to break it" - and watch *how* they wield the tooling. Because wielding tools, not reciting from memory, is the real competency now. ## Not animals, but ghosts Finally, a metaphor that organizes the whole mindset. Karpathy says we're not building animals - we're *summoning ghosts*. > "We're not building animals. We are summoning ghosts." These are statistical simulation circuits - a pre-training substrate with RL bolted on - not intelligences shaped by curiosity and evolution. Yelling at them does nothing. The value of the metaphor is practical: it's a *mindset*. Be appropriately suspicious and check empirically what works, instead of assuming the model "understands" the way a human does. And here I come back to you. **Which paradigm is your product in** - 1.0, 2.0, or 3.0? And when the big LLMs can do your task directly, **what will remain your moat?** Karpathy leaves four answers: 1. **Data** - proprietary datasets, your own RL environments, a model trained on *your* knowledge and style. 2. **Prompts and context** - the context systems and knowledge base you've built to steer the engine. 3. **System design** - sensors, actuators, UX, and verification loops around the model. The engine is just the engine; you design the car. 4. **Trust** - brand and accountability for the outcome. "I trained an LLM on all of Elon's knowledge" sounds one way from a random person, another way from Elon himself. The model can't buy that. If your edge doesn't sit in one of those four moats, it's worth asking yourself that uncomfortable question sooner rather than later.

Want to rebuild your product around agentic engineering?

I'll help you assess which paradigm you're in, where your moat is, and how to deploy agents without dropping the quality bar - from context architecture to verification loops.

Book a free consultation
## Useful Resources - [Andrej Karpathy: From Vibe Coding to Agentic Engineering](https://www.youtube.com/watch?v=96jN2OCOfLs) - Sequoia Capital, 29:49. The primary source for the whole thesis: vibe→agentic, verifiability, jagged intelligence. - [Kaparthy revealed the most profitable business to build (Software 3.0)](https://www.youtube.com/watch?v=hJNp9RwK-Uw) - Dream Labs AI, 14 min. The business angle on the paradigms and the four moats. - [Vibe coding: a guide](/en/blog/vibe-coding-guide) - how to raise the floor, the other side of this distinction. - [A Karpathy-style knowledge base](/en/blog/karpathy-llm-wiki-knowledge-base) - how to build the data and context moat. ## FAQ
### How does Software 3.0 differ from Software 1.0 and 2.0 according to Karpathy? Software 1.0 is writing explicit code rule by rule, Software 2.0 is programming by curating data and training neural-network weights, and Software 3.0 is **prompting** - steering the LLM through the context window. In 3.0, what you put into the context window becomes your lever over the model treated as an interpreter. The practical consequence: instead of building a tool, you design a system around an LLM engine that delivers a finished outcome.
### What's the difference between vibe coding and agentic engineering? Vibe coding "raises the floor" - it lets anyone build software, which is democratizing. Agentic engineering "preserves the ceiling" - it's an engineering discipline focused on keeping the professional quality bar (no new vulnerabilities, accountability for the code) while accelerating with the help of agents. In short: vibe coding is accessibility, agentic engineering is quality under the pressure of speed.
### What is jagged intelligence in AI models? Jagged intelligence is the phenomenon where a model has peaks of capability in verifiable domains, like code or math, and valleys in other tasks. It stems from training in RL environments with verification rewards - the model is excellent where the result can be checked automatically. Hence the paradox: the same model will refactor 100,000 lines of code yet stumble on a simple everyday task.
### What can't you outsource to AI agents when building a product? You can't outsource understanding, judgment, and oversight - as Karpathy puts it, "you can outsource your thinking, but you can't outsource your understanding." The human stays the owner of the specification, the plan, and the top-level design, while the agent fills in the blanks. Fundamentals matter too (understanding what happens underneath), because an agent can make a nonsensical mistake it won't catch on its own.
### How do you build an "agent-native" product, for agents and not just humans? Describe systems to agents first (e.g. a block of instructions to paste, not human-oriented documentation), decompose work into sensors that read the world and actuators that act on it, and keep data legible to the LLM. The goal is a world where "my agent talks to your agent." It's also worth refactoring hiring: instead of puzzles, give a real project and watch how the candidate wields the tooling.
--- # Pre-revenue startup without a finance department. AI agent system anatomy. Source: https://pawel.lipowczan.pl/en/blog/ai-agent-system-skills-rules-shared-context Published: 2026-05-09 Today is May 9, 2026. Two days ago I gave a talk at NoCode Poland #4 - *"How 2 founders do the work of a whole team"*. The title sounds like clickbait, but a week earlier, on May 1, part of what I was talking about **didn't exist yet**. On May 1 at 200IQ LABS - the pre-revenue company I run with Przemek - there was no finance department. Zero April management report. Zero category-level cost control. Zero spec. **On May 3 we did the first April close**: 30+ classification rules trained from scratch, a management report in `monthly/2026-04.md`, and a real EBITDA of −16,804 PLN documented with three plan-correction decisions for the coming months. Two days from zero to a working management system. Honestly - I did this partly for the talk. **Event Driven Development™** (the event was the talk). But this system was overdue. I'd rather ship PRs to production than sit on numbers - that's the founder truth. The talk just forced the timing. This article isn't about *"AI replacing accountants"*. It won't - formal accounting via inFakt + an accountant handles compliance, and still does. This article is about **how to architecturally build an AI agent system that genuinely replaces specialist roles in the management layer** - where the department doesn't exist because the company is too small to hire and too big to improvise. Three pillars I keep returning to throughout the text: 1. **Skills automate processes.** Workflow becomes a noun - `/finances close 2026-04` instead of 6 manual steps. 2. **Determinism through rules.** Hybrid rules-first + LLM fallback + learning loop. The system converges from probabilistic to deterministic. 3. **Shared context.** CLAUDE.md, MEMORY.md, the `context/` structure. Without it, subagents are isolated chats. With it - a system that remembers the company. ## Architecture: one diagram, three layers ![AI agent system architecture at 200IQ LABS - main agent (Claude Code orchestrator) → 5 subagents (CFO, Marketing, Legal, Tax Advisor, Business Consultant) → tools layer (Python tools + External MCP) → shared context (CLAUDE.md, MEMORY.md, context/)](/images/architecture-system-agentow-ai.webp) This diagram is the spine of the article. Each of the three layers maps to one of the pillars I'll come back to. **The main agent** is Claude Code acting as an orchestrator. From its perspective I'm the user; it's the entry point. Routing to subagents happens through skills - I type `/finances close 2026-04`, the agent loads the `finances` skill, persona switches to CFO. **Specialized subagents** - CFO, Marketing, Legal, Tax Advisor, Business Consultant - aren't separate processes or separate models. They're **combinations of skill + persona + scoped context**. CFO reads `context/finances/`, Marketing reads `context/qamera/` (because Qamera is our product) and `context/brand/`, Legal reads `context/legal-entities.md` and `context/company.md`. Each role only sees what it needs - context hygiene. **The tools layer** is a hybrid: Python (default for planned workflow) + External MCP (for ad-hoc or when an API forces it). Stripe, Revolut, Airtable are Python wrappers in the repo - because we know up-front how we use them. inFakt is MCP because the costs endpoint returns 403 when an accounting office handles the entity - the official MCP server bypasses that limitation. Qamera AI has its own MCP, used by both external agents and ourselves - because for ad-hoc exploration of our own product, MCP is just faster than writing a script per query. **Shared context** is the substrate of the entire system. CLAUDE.md loads into every conversation - project instructions, conventions, paths. MEMORY.md is persistent across sessions. `context/` is a structured knowledge base: finances, clients, projects, operations. Subagents read from this base - and that's why they aren't isolated chats. Now in order. ## Pillar 1: Skills automate processes **A skill is a unit that encloses a process** - trigger, steps, decisions, integrations. Before skills, your workflow lives in your head and someone's notebook. With skills, the workflow becomes a noun: `/finances close YYYY-MM`, `/ingest`, `/slides:new`. Command instead of memory. Sounds banal. It isn't. Without a skill I improvise every time: *"hey, pull Stripe data for April, then Revolut, classify transactions, check accruals, write the report"*. That means I reinvent the same sequence 6 times a day. With `/finances close 2026-04` - one command, deterministic phase sequence. ### `/finances close` - 6 phases ![Six phases of monthly close in the `/finances` skill: PULL → CLASSIFY → REVIEW → ACCRUALS CHECK → COMMIT → REGENERATE. PHASE 3 and 5 are human-in-the-loop](/images/diagram-close-process.webp) ```text PHASE 1: PULL (~2 min) Revolut + Stripe (Python) + inFakt (MCP) + tech-stack PHASE 2: CLASSIFY (~30s) rules.yaml deterministic → LLM fallback PHASE 3: REVIEW (~10-15min interactive) [a/c/r/s] learning loop PHASE 4: ACCRUALS CHECK (~2 min) accrued-liabilities.yaml matching PHASE 5: COMMIT (~10-20 min) mandatory narrative blocks close PHASE 6: REGENERATE (~30s) 3 dashboards, AUTO:START/END markers ``` Each phase is **idempotent** - restart after a crash doesn't duplicate data. You can run this close, interrupt after PHASE 3, come back an hour later, and continue from where you stopped. This isn't pipe-and-pray; it's a transactional operation in the sense of "you can repeat it and nothing bad happens". ### Video: one close, live

Full April 2026 close run as /finances close 2026-04. ~7 minutes, 6 phases, 2 QA iterations stayed in the recording - meta-message: you can correct things by talking to the agent.

The recording keeps **two QA iterations** - moments where I stopped the agent, noticed something was off, and asked for a correction without restarting the whole phase. I deliberately didn't cut them out. That shows something that gets lost in marketing demos: a workflow with an agent isn't a single-pass script. You can correct via conversation, the agent adjusts state, continues. ### Generalization: any process can be a skill `/finances close` is the flagship skill, but the same mechanic works everywhere. `/ingest` classifies files dropped into `inbox/` - I described it in detail in the [PIT-38 case study](/en/blog/polish-pit-38-claude-code-case-study). `/slides:new` generates a presentation skeleton. Meta-twist: the NCP4 slides I mention in the intro were created with the same workflow. **The very `/blog-article-writer` skill used for this article** has its own phases: prime → plan → execute → validate → translate. ### Mandatory narrative blocks close PHASE 5 requires that a human writes three narrative sections in `monthly/.md`: 1. What happened unexpectedly (vs plan). 2. Decisions taken during the month. 3. Plan for next month / plan correction. **Empty narrative blocks the close.** The system won't close the month until you write something in each section. This looks like pointless formalism - and from the *"I just want to close fast"* perspective, it is. But from the *"six months later I'm staring at EBITDA variance −3,160 vs plan and don't know why"* perspective, it's the only way decisions retain durable memory. **Reflection is part of the process, not optional.** This pattern is broader than finances: every time you automate something where value lies in **understanding** rather than **execution** - add a human-in-the-loop checkpoint that can't be skipped. Otherwise you'll automate execution and destroy understanding. ## Pillar 2: Determinism through rules Transaction classification is a classic cold-start problem. Pure rules: every new vendor needs manual setup, a month with no new vendors is a good week, a month with 5 new ones is an evening. Pure LLM: non-deterministic, cheap at scale, but unacceptable in finance - audit, regulator, future inspection. The third path is hybrid: **rules-first + LLM fallback + learning loop**. ### Rules-first ```yaml # rules.yaml - fragment - id: gworkspace-via-gcp pattern: memo: "GCPLD" source: infakt amount_range: [200, 500] classify: category: opex/saas unit: null note: "Google Workspace billowane przez Google Cloud Poland (300-330 PLN/mc)" created_at: 2026-05-04 - id: gcp-infakt pattern: memo: "GCPLD" source: infakt amount_range: [500, 999999] classify: category: cogs/ai-generation unit: qamera note: "Faktury GCP w inFakt z prefiksem GCPLD i kwotą >500 PLN = compute" created_at: 2026-05-04 ``` The strongest detail: **two rules for the same vendor** (`GCPLD` in inFakt), differing only in `amount_range`. 200-500 PLN → Google Workspace (category `opex/saas`). >500 PLN → compute GCP for the Qamera pipeline (category `cogs/ai-generation`). This is the type of nuance an LLM gets **differently every time**. One day it classifies all `GCPLD` as Workspace, the next as compute, the third randomly. Here we have the answer encoded as code. Auditable, deterministic, idempotent. The `note:` field matters as much as the pattern. Without that note, six months later I look at the rule and don't know why splitting by `amount_range` makes sense. **A rule without a note is long-term debt.** ### LLM fallback with few-shot examples When no rule matches, the agent falls back to LLM with a few-shot prompt. Each example in `examples.yaml` has explicit `reasoning`: ```yaml # examples.yaml - fragment - transaction: date: 2026-04-10 amount: -127.40 memo: "BYTEPLUS API USAGE PAY-AS-YOU-GO" source: revolut classification: category: cogs/ai-generation unit: qamera reasoning: "Byteplus to provider Kling AI używany do generacji video w produkcie" ``` The LLM gets 15-25 such examples per category. It returns a classification plus its own `reasoning`. **The `reasoning` is visible to the user** in the REVIEW phase - it's not a black box. I see not just "category X" but why. ### Learning loop: REVIEW [a]ccept / [c]hange / [r]ule / [s]kip A loop per LLM-classified transaction. Four decisions: | Option | What it does | |---|---| | `[a]ccept` | LLM was right, but it's a one-off case | | `[c]hange` | LLM got it wrong, manual correction | | `[r]ule` | Accept + create a rule for similar cases | | `[s]kip` | Don't know, we'll come back next close | **Golden rule:** if the transaction is likely to repeat → `[r]ule`. 30 seconds now, you save minutes over the coming months. Every `[r]ule` freezes the LLM's guess into a deterministic rule in `rules.yaml`. Next close - the same transaction lands in "auto-classified" deterministically. **Real number:** the first close (April 2026) generated **30+ rules from zero**. Three pattern types surfaced immediately: 1. **By memo** - pure pattern matching. Byteplus, Cursor, Meta Pay. 2. **By inFakt invoice prefix + amount** - `GCPLD` with `amount_range: [200, 500]` is Workspace, with `[500, ∞]` is compute. Same vendor, two categories. 3. **By reseller pattern** - Paddle 24 EUR ≈ 100 PLN is n8n cloud. But Paddle with a different amount = REVIEW (could be a different product sold via Paddle). The system converges from probabilistic to deterministic. This is a rare design - most budget classification tools are either 100% deterministic (rigid rules, lots of manual work) or 100% probabilistic (LLM/ML, no audit). The hybrid with `[r]ule` as a bridge between the two gives you LLM cold-start and long-term audit of deterministic rules. ### Idempotent regeneration as an orthogonal determinism mechanism The second layer of determinism is the dashboards. PHASE 6 regenerates 3 files: `_dashboard.md`, `_alerts.md`, `_runway.md`. Question: what about manual notes I added during the month? If regeneration overwrites the whole file - my notes vanish and I stop writing notes. Solution: scope markers. ```markdown # 200IQ LABS - Finances Dashboard ## Cash position & runway | Metric | Value | |---|---| | Cumulative shareholder loans | 74,300 PLN | | Confirmed financing (June) | +100,000 PLN | | Accrued liabilities (UoD payout May) | −24,000 PLN | | Run-rate burn | ~14-14.5k PLN/mo | ## Manual notes (outside markers - survive regeneration) - 2026-05-04: first system close. Manual dashboard regeneration. - Plan finalize: not done - `draft` mode kept deliberately until next close. ``` The `AUTO:START` / `AUTO:END` markers limit the scope of regeneration. Everything outside the markers **survives every regeneration**. Without it, nobody writes manual notes in auto-generated documents - and those documents become dead. This simple pattern decides whether dashboards get used or ignored. ## Pillar 3: Shared context (the substrate) The third pillar is the layer that makes all subagents coherent. Without shared context, CFO and Marketing don't know about each other - you have 5 isolated chats, each with its own version of truth. With shared context, you have **a system that remembers the company**. ### Three layers of shared context - **CLAUDE.md** (per-project) - instructions, conventions, paths, "what should the agent do when in doubt". Loads into every conversation. - **MEMORY.md** (cross-conversation) - auto-memory, persistent across sessions. This is where *"the user prefers short summaries"*, *"project X uses OpenSpec"*, *"today's date is YYYY-MM-DD"* live. - **`context/`** (knowledge base) - a structured base per domain. `context/finances/`, `context/qamera/`, `context/clients/`, `context/operations/`. Directory layout (simplified): ```text agentic-ai-system/ ├── CLAUDE.md # project instructions ├── memory/ │ └── MEMORY.md # auto-memory, persistent ├── context/ │ ├── finances/ # ← read by skill CFO │ ├── qamera/ # ← read by skill Marketing (revenue + brand) │ ├── clients/ # ← read by skill Business Consultant │ └── operations/ ├── skills/ │ ├── finances/SKILL.md │ ├── ingest/SKILL.md │ └── slides/SKILL.md └── tools/ ├── stripe/ # Python wrappers ├── revolut/ # Python wrappers └── airtable/ # Python wrappers ``` Each subagent has a scope - it doesn't read everything. CFO reaches into `context/finances/` and `context/operations/`, with read-only access to `context/qamera/` for revenue. It doesn't read `context/clients/` (that's the Business Consultant's scope) because it doesn't need to. **Context hygiene** is the second multiplier of value - without it, every CFO question pulls half the company into tokens. ### Python tools (default) vs External MCP (when forced) > **Rule: Python script for planned workflow. MCP for ad-hoc or when the API forces it.** | Integration | Approach | Why | |---|---|---| | Stripe (Qamera subscriptions) | Python | We know up-front what we need - wrapper is part of the spec | | Revolut (company transactions) | Python | Fixed sequence in `/finances close`, own wrapper | | Airtable (CRM) | Python | Repeatable CRM operations, own wrapper | | **inFakt** (accounting) | **MCP** | Costs endpoint returns 403 when an accounting office handles the entity - MCP **forced** | | **Qamera AI** (our product) | **MCP** | Both us and external agents. REST API is also available, but for ad-hoc exploration MCP is just faster to wire up | **Why Python > MCP, when the workflow is planned:** 1. **Token economy.** MCP loads tool descriptions into the context of **every** conversation. Script = zero overhead. The script is invoked only when the agent decides it's needed. 2. **Determinism.** A script does exactly what it says. No "the model interpreted the tool differently in another conversation". 3. **Trivial to create.** *"Agent, I need a script that pulls all active Stripe subscriptions on day X"* → agent generates, I code-review, commit. 15 minutes. 4. **Full control.** Code in the repo, versioned, readable for another developer. I can debug in Python, add logging, change output format - without negotiating with the MCP protocol. **Why MCP, when we use it ad-hoc:** If I'm still **discovering** how I'll use something - writing a script is premature. I don't know yet which fields I care about, which filters, what output format. MCP gives the agent immediate access to the API; after several iterations I see the pattern worth freezing - and then I write the script. **The script is part of the workflow specification; MCP is a tool for exploration.** That's exactly how we use Qamera: external agents hit the MCP, we do too - when we want to quickly check something in our own product. If we had a steady workflow like *"generate X for client Y every week"* - we'd write a script. Right now we don't, so MCP suffices. With inFakt we have no choice - the REST API returns 403 on the costs endpoint when an accounting office handles the entity. Here MCP isn't an architectural choice, just the only available path to data. ### Subagents read only their scope Back to the first diagram: each subagent (skill + persona) has a scope on its context line. The `finances` skill has front-matter that explicitly declares: ```yaml context_scope: - context/finances/ - context/operations/ - context/qamera/ # read-only, for subscription revenue ``` Without this restriction, the agent asks *"what were the transactions in March"* and pulls the whole company's context into tokens. With the restriction - it only loads `context/finances/transactions/2026-03.yaml` plus rules. Fewer tokens, faster response, lower hallucination risk. I wrote about the same pattern in a different context - the `inbox/` / `archive/` / `data/` separation in the [PIT-38 case study](/en/blog/polish-pit-38-claude-code-case-study). The pattern generalizes: **the agent works on clean, processed files, not raw data**. ## Case study: real numbers from April 2026 Proof that this isn't a demo. The first close, executed on 2026-05-04, generated: | Line | Actual | Plan | Δ | |---|---:|---:|---:| | Revenue | 347 PLN | 598 PLN | −42% | | Costs | 17,151 PLN | 14,242 PLN | +20% | | EBITDA | −16,804 PLN | −13,644 PLN | −23% | | `cogs/ai-generation` | 1,709 | 700 | +144% 🔴 | | `opex/marketing` | 2,996 | 1,500 | +100% 🔴 | | `opex/saas` | 2,112 | 1,633 | +29% 🟡 | **Three decisions arising from the close** (recorded in `monthly/2026-04.md`): 1. **GCP +144% vs plan** → the plan was wrong (700 PLN), the real run-rate is historically 1,000-1,600 PLN/mo. Plan May-Dec raised retroactively to 1,500-1,800 PLN. Conscious decision **NOT to optimize GCP** - part of this is content marketing (own + materials for clients). The "this is actually marketing" classification doesn't change the category (GCP stays in `cogs/ai-generation`), but it justifies a higher plan. 2. **Cursor billed 4×/mo** (716 PLN vs plan 217) → planned **Cursor → Claude** migration (rigid ~90 EUR/mo cap, predictable cost instead of usage-based). 3. **Meta Ads off from May** - strategy was wrong, direction TBD. Control plan: in the May close we'll verify spend < 300 PLN. **The point:** the decisions follow **from the system, but the system doesn't make them**. The system shows variance and requires a narrative in PHASE 5. The decision is mine - but I have it documented for the moment six months later when I forget why. Break-even target: September 2026 (revenue 12,184 vs costs 12,842, EBITDA −658). Aggressive plan - needs validation in upcoming closes. **Plan = a living document**, corrected after every close as run-rate shifts. Drift > 30% = revision required. ## What I don't recommend A risk map, mirroring the section from the PIT-38 case study. - **I don't recommend this setup for someone without programming comfort.** Markdown, YAML, git, terminal, editing convention files - requires fluency with tools. Without it, the time lost on setup eats all the gains. Safer path: a ready-made controlling SaaS (Causal, Runway, Pry) + Excel + an accountant. - **Manual first, automate after pain.** The first 2-3 closes MUST be manual. Without a manual run, half the rules would be designed wrong. The first close revealed 30+ patterns to encode - including non-obvious ones (e.g. that `GCPLD` is two categories, not one). **Scripts come after, not before.** - **A management system ≠ formal accounting.** inFakt + an accountant do compliance. This system makes decisions. Don't substitute one for the other - they complement, not replace. I still pay my accountant a monthly retainer; this system doesn't remove that line from the budget. - **LLMs make arithmetic errors.** Send sums to a compute engine (Excel, calculator, MF system). LLMs add value in **structure and interpretation**, not arithmetic. I wrote about this in the [PIT-38 case study](/en/blog/polish-pit-38-claude-code-case-study) - I left 6 cents on the table because the LLM tripped subtracting two six-digit numbers. - **It doesn't scale linearly beyond one company.** The system is designed for 200IQ LABS. The patterns generalize (skills, rules, context); the specific YAMLs don't. *"Don't copy the setup, extract the principles"* - that was the motto of the NCP4 talk and it's the motto of this article. ## Key takeaways 1. **Skills make a process a noun.** `/finances close YYYY-MM` instead of 6 manual steps. Every repeatable workflow deserves its own skill - with phases, idempotency, and mandatory checkpoints where value lies in understanding. 2. **Determinism is built iteratively.** Rules-first + LLM fallback + the `[r]ule` option in REVIEW = the system converges from probabilistic to deterministic. After the first close we had 30+ rules. After the tenth we'll have ~150, with the LLM invoked rarely. 3. **Shared context (CLAUDE.md, MEMORY.md, `context/`) is substrate, not decoration.** Without it, subagents are isolated chats. With it - a system that remembers the company and shares knowledge across roles. 4. **Python script for planned workflow, MCP for ad-hoc.** A script is part of the spec - token-efficient, deterministic, code-reviewable. MCP gives immediate access when you don't yet know how you'll use something, or when the API forces it. 5. **Reflection is part of the process, not optional.** Mandatory narrative blocks close. Without it, six months later you don't know why something was. I'd rather ship code than count beans - and the agents help me with both. This system was overdue. The driving event was the talk. It works.

Running a pre-revenue startup and thinking about an AI management system?

I help founders and technical consultants design AI agent architectures for company operations - skills, deterministic rules, shared context. I'll show you what such a setup could look like for you.

Book a free consultation
## Useful Resources - [Claude Code - documentation](https://docs.claude.com/en/docs/claude-code/overview) - reference for skills, settings, MCP, hooks - [Model Context Protocol - spec](https://modelcontextprotocol.io/) - what MCP is and when it makes sense - [PIT-38 case study](/en/blog/polish-pit-38-claude-code-case-study) - the same architectural pattern in a tax-filing context - [Skills 2.0 - multi-agent system for company management](/en/blog/skills-2-0-multi-agent-system-company-management) - broader Skills 2.0 context and the role of subagents - [OpenSpec workflow - structured AI work](/en/blog/opsx-workflow-structured-ai-work) - spec-driven development as the foundation for skills - [Spec-driven SEO on the portfolio and Qamera AI](/en/blog/spec-driven-seo-portfolio-qamera-ai-case-study) - another case study, the same type of workflow ## FAQ
### Does this pattern (skills + rules + shared context) work only in Claude Code, or in other agents too? The pattern is agent-agnostic. Skills map to commands/functions in other systems (Cursor commands, n8n workflows, custom CLIs). Rules + few-shot examples are a standard ML technique. Shared context (markdown + directory structure) works anywhere the agent has file access. **The specific mechanisms** (`CLAUDE.md`, MCP, `MEMORY.md`) are Anthropic-specific, but **the architectural principles** generalize to any LLM-driven system.
### How long does it take to build a system like this from scratch? The first working close - with OpenSpec, YAML schemas, and 30+ rules - took us two days of intensive two-person work. Full stabilization (auto-pull scripts, scheduled close, integration with formal accounting) is planned over 4-6 weeks at a regular monthly cadence. **Every day you use a manually designed skill is a productive day** - you're not waiting for the system to be complete, because value shows up immediately after the first close.
### Why a Python script instead of MCP, if MCP is the standard? MCP loads tool descriptions into the context of **every** conversation - that's a token cost and a potential interpretation ambiguity across sessions. A Python script is zero overhead, deterministic output, code in the repo for code-review. **We write scripts for planned workflow** (Stripe, Revolut, Airtable - we know up-front what we need, so the script is part of the spec). **We use MCP for ad-hoc** (exploring our own Qamera, when we don't yet know which queries will repeat) **or when the API forces it** (inFakt returns 403 on the costs endpoint when an accounting office handles the entity). Pattern: after several ad-hoc iterations, you can see what's worth freezing into a script.
### Won't the LLM make mistakes in financial classification? What about audit for the regulator? We design the hybrid (rules-first + LLM fallback) specifically to minimize LLM-classified transactions. Every LLM-classified transaction goes through REVIEW - a human sees the `reasoning` and decides `[a]ccept`/`[c]hange`/`[r]ule`/`[s]kip`. After the first close we had 30+ rules that deterministically catch 80%+ of subsequent transactions. **The audit trail lives in `rules.yaml` plus `monthly/.md`** - every decision documented, every rule has a `note:` with justification. Formal accounting (compliance) still goes through inFakt + an accountant; this system is managerial, not compliance.
### What does mandatory narrative do if I have an emergency close and don't have time to write 3 sections? The three sections (surprises / decisions / plan correction) block the close until filled - that's a deliberate decision. In an extreme case you can write a single sentence in each section (*"no surprises vs plan"*, *"no decisions"*, *"continuing without corrections"*). The system doesn't grade narrative quality, only its presence. **The point is that six months later, looking at EBITDA variance, you need any context - short is better than none.** This formalism is intentional.
### Does this system work for a revenue-stage company, not just pre-revenue? Yes, the pattern scales - P&L categories and caps need adjusting, mandatory narrative becomes even more important (more transactions = more decisions to document), the learning loop yields better ROI (more repeating vendors). **But in a post-revenue company with high volume you probably need a full ERP** - a management system in markdown and YAML hits a ceiling at ~hundreds of transactions/month and a few decision-makers. **Our scope:** pre-revenue → early-revenue, 1-5 people in the company. Above that scale the architecture needs harder tools.
--- # 3 days to PIT-38, no accountant. 2 hours were enough. Source: https://pawel.lipowczan.pl/en/blog/polish-pit-38-claude-code-case-study Published: 2026-04-29 Today is 29 April 2026. The PIT-38 filing deadline expires tomorrow. The day before yesterday in the evening I started with a single sentence - _"create a new PIT-38 project"_. Yesterday afternoon the return was filed with the Polish tax office, the tax was paid, and the official receipt (UPO) was sitting in `output/`. Total **~2 hours** of active work. **5,430 transactions** across two crypto exchanges. **174,895.50 PLN** of unused crypto cost basis carried forward from 2024. **2 PIT-8C forms** from brokers, **1 supplementary report** on foreign dividends. Filed and accepted by the tax office **51.5 hours** before the deadline. Final tax due: **172 PLN**. My accountant usually handles this. This year I missed the timing window - and that turned out to be one of the more instructive experiences I've had with an agentic workflow. This article is not about how I learned to file a PIT-38. I knew what to do, because I've been doing it for years (through my accountant). It's about how **two hours of active work with a properly configured agent** reproduced the work of an expert service - with better documentation than I usually got from my accountant. > Quick context for non-Polish readers: **PIT-38** is the Polish annual tax form for capital gains income - stocks, mutual funds, foreign dividends, and crypto. It's notoriously fiddly because of multi-source income, foreign currency conversion via central bank D-1 rates, and a multi-year cost-basis carry-forward mechanism for crypto. Filing deadline is 30 April for the previous tax year. ## My accountant usually files my PIT-38 I file a PIT-38 every year. Stocks, funds, crypto, foreign dividends. Never alone. Always through my accountant - I send her the documents, she fills out the form, double-checks, files. It works. This year I missed the timing. Deadline 30 April, and I started gathering documents on 27 April in the evening. No accountant will take a project with a 3-day buffer - and rightly so. I was forced to do it myself. Two paths: (a) panic and a hasty last-minute filing, or (b) test whether the AI workflow I use to build things for clients can replace my accountant in two days. I picked (b). It worked. That changes the article's stakes. This is **not** a "lifehack for people who don't want to hire an accountant." It's a **case study of when AI automation realistically reaches into expert-level service work** that until now was handled by humans. ## Project architecture: every directory has one responsibility I started with the prompt _"create a new PIT-38 project"_. No brief, no requirements doc, no task list. The agent asked about the tax year and deadline, reached for the internal `_template.md` convention (which it knows from the repo's `CLAUDE.md`), and laid out: ```text PIT_38/ inbox/ # drop zone, raw files archive/ # post-processing (agent does NOT read) data/ # knowledge - agent reads freely deliverables/ # checklist, blog inputs output/ # final return PDF, UPO project.md # goal, status, decisions catalog.md # file index ``` Every directory has one responsibility. The agent does not look inside `archive/` or `output/` unless I explicitly ask it to. This is **not security** - it's **context hygiene**. The agent works on clean, processed `data/*.md` files, not on 4,096 rows of raw CSV with every question. It sounds banal. It isn't. Without that separation, every question like _"what were the March transactions?"_ would pull the entire raw CSV into context - either blowing the token budget or forcing the model to "sum on its fingers." With separation: the agent gets a finished `data/crypto-summary-2025.md` with 11 line items and works on structure, not on raw input. That's the first multiplier of value - and it works **only because** the convention is described in `CLAUDE.md`. Without that file, this project would have taken as long as doing it manually. ![Workflow funnel: 5 independent data sources converge into /ingest, the agent classifies 5,430 raw transactions into 11 taxable events, the return is filed with the Polish tax office with a 174,895 PLN cost-basis buffer rolled forward to 162,948 PLN](/images/diagram-pit-38-funnel-en.webp) ## `/ingest`: one command, four steps I dropped 7 files into `inbox/`: two PIT-8C forms (XTB and SFIO), three CSVs from crypto exchanges, the dividend report, and last year's filed return as historical reference. Command: `/ingest PIT_38`. What happens under the hood: 1. **File type identification** - the agent recognises a PIT-8C from its header structure, a crypto report from typical column patterns (timestamp, asset, type, amount). 2. **Routing** - PIT-8C goes to section C of the return, crypto reports to section E, dividends to section G. Each file gets its own `data/{source}-{kind}-2025.md` with normalisation. 3. **Extraction** - key items (revenue, cost, withholding tax) are pulled from PDFs, transaction types are classified from CSVs. 4. **Archiving** - raw files move from `inbox/` to `archive/`, `catalog.md` is updated, `project.md` gets a changelog entry. A concrete example: PIT-8C from both XTB and SFIO ended up summed in `pit38-calculation.md` in section C, using line numbers consistent with the **PIT-38 form version (18)** for tax year 2025. The agent itself noticed that this was not the same form layout as 2024 - more on that shortly. The lesson at this stage: `/ingest` is **not batch processing**. The value lives in the iterative loop, which I'll show next. ## Six things I didn't expect from the agent This is the heart of the article. Six concrete things the agent did across those two days that, had I worked alone with Excel, I either wouldn't have done at all or would have done with errors. ### 1. The 174,895.50 PLN buffer - auto-pulled from PIT-38 2024 Under Polish PIT law, unused crypto cost basis rolls forward indefinitely (Art. 22 §16 of the PIT Act). It's a standard mechanism for active crypto investors. I knew this buffer existed - I've been filing PIT-38 for years. But I knew it from the perspective of _"tell my accountant about it,"_ not from the perspective of _"enter it in field 38 of the new form precisely."_ On the first ingest pass, the agent asked for last year's return. I uploaded the 2024 PIT-38 PDF. The agent itself extracted the amount from the relevant line and entered it in the matching slot of the new return - where it becomes "costs from prior years." Value to me: that buffer is usually _"remembered"_ by the accountant, who has my prior year's return in her CRM. Without her, I would have had to remember on my own to download the PDF, open it, find line 38. **The agent did that with a single question.** This is an underappreciated class of value - automating "easily accessible memory of prior years' context." Accountants do it. Agents with access to historical files in `archive/` do it. ### 2. PIT-38 form version (17) vs (18) - the small detail that breaks a return In PIT-38 for 2024, prior-year crypto costs lived in one line. In 2025, that line moved. Section C in version (18) gained an extra row for exemptions under Art. 21 §1.105a - and everything below it shifted by two lines. Many blog posts and online forums still reference the **old line numbering**. If I had entered the data into the old positions, the return would either be rejected outright or, worse, accepted with a calculation error that would surface during an audit. The agent caught this after I uploaded the blank PIT-38(18) form template as reference. It compared the structure with last year's filed return, flagged the differences in `data/sources.md`, and used the new numbering in the final calculation. Lesson: **trust reference documents, not the model's memory.** A current form template uploaded > whatever the model picked up from training data a year ago. ### 3. 5,430 transactions → 11 taxable events 4,096 transactions from one platform plus 1,334 from another. 5,430 raw rows in total. After classification: **11 taxable events** that go on the return. The rest are platform-internal operations - transfers, technical adjustments, accruals, small reconciliations - that don't need to be reported. The LLM's value here is not in **summing**. It's in **classification**. Each of the ~25 transaction types in the report was assigned to one of three categories: taxable event / internal operation / neutral. Each classification was backed by a documented legal rationale in `data/sources.md`. > I won't go into specific transaction types or legal interpretations - that's beyond the scope of this case study and requires individual consultation with a tax advisor. I'm only showing the **scale** of the classification work. Without that classification, trying to file 5,430 raw CSV rows manually would have been either impossible in two days or would have led to a dramatic overstatement of the tax base - if I had treated every transaction as a disposal. The scale that the LLM reduces here is non-trivial. ### 4. Central bank D-1 FX rates - 10 edge cases with a holiday calendar Every transaction in EUR/USD has to be converted to PLN using the **mid-market rate from the National Bank of Poland (NBP) on the business day before the disposal date** (Art. 11a of the PIT Act). Sounds simple. In practice 10 of the 11 transactions needed individual handling, because they fell on non-business days: | Transaction date | Day of week | NBP rate from | | --- | --- | --- | | 2025-02-15 | Saturday | Friday 2025-02-14 | | 2025-03-16 | Sunday | Friday 2025-03-14 | | 2025-05-02 | Friday | Wednesday 2025-04-30 (1 May = public holiday) | | 2025-05-18 | Sunday | Friday 2025-05-16 | | 2025-06-15 | Sunday | Friday 2025-06-13 | The agent did this manually for each of the 10 transactions: computed the day of week, checked national holidays (1 May, 3 May, Corpus Christi, 15 August, 1 and 11 November, 25-26 December), pulled the rate from NBP, applied the conversion. This is **exactly the kind of work where a human brain trips up quickly** - every edge case requires checking a calendar separately. A human would make 1-2 mistakes per 10 transactions (forgetting 1 May is a common Polish-tax bug). The agent made 0. The class of task _"many small mechanical decisions with subtle rules"_ is an LLM's strength. Not summing. Not creativity. Systematic application of rules to each of 10 situations without boredom and without omission. ### 5. The 6 groszy the LLM couldn't add up My calculation for section C: a loss of **966.98 PLN**. The official tax-office portal after data ingest: **966.92 PLN**. A 6-grosz (0.06 PLN) discrepancy. Arithmetic error from the LLM when subtracting two 6-digit numbers (16,188.05 − 15,221.13). Lesson: **LLMs make arithmetic errors on 6-digit additions.** Leave summing to a calculator, Excel, or the tax-office system. The LLM has value in **structure and interpretation**, not arithmetic. This isn't a failure - it's a map of where to use the tool and where not to. In this specific case the tax-office portal corrected my error for free. In another scenario - say, a paper filing - a 6-grosz delta would have gone to the tax office and probably no one would have cared, but the principle stands: **arithmetic to a compute engine, not to a language model**. ### 6. Ingest as progressive discovery - the strongest beat The most valuable thing the agent did across those two days was not summing 5,000 transactions. It was **asking about things that weren't on my list**. I didn't have a list of _"what's needed for PIT-38."_ I uploaded what I had at hand. After each iteration, the agent said: _"OK, this is here, but X is missing"_ or _"watch out, this implies Y, check Z."_ The key moment: **foreign dividends**. - I didn't know I had dividends from 2025 (small ETFs in XTB, around 958 PLN gross). - I didn't know they have to be reported separately from everything else (section G of PIT-38, Art. 30a, 19% PL minus withholding). - After the first ingest, the agent asked: _"and what about dividends? PIT-8C doesn't include them, XTB issues a separate report."_ - I downloaded the **XTB Supplementary Report for PIT-38** - and there were 182 PLN of 19% PL tax on foreign dividends to declare. Without that question, I would have filed the return **without section G**. Result: a 172 PLN underpayment, audit risk, late-payment interest. The second moment: **2024 history**. On the first ingest the agent asked whether I had last year's return. I uploaded the PDF - and that's when the 174,895.50 PLN buffer described above surfaced. Without that move I would have paid ~2,200 PLN in tax instead of 0. This is **the essence of the value**: not _"the agent summed transactions,"_ but **"the agent knew what I didn't know."** Iterative questioning + classifying every document against the PIT-38 taxonomy = impossible to achieve manually without specialised tax knowledge. Ingest workflow ≠ batch processing. The value is in the loop: you upload → the agent classifies → it identifies gaps → it asks for more → you upload again. After 3-4 iterations you have a complete picture you couldn't have assembled on your own. ## A short primer: how crypto taxation works in Poland A small explainer for readers outside Polish crypto-tax topics, so that "buffer of 174k" doesn't read as "Pawel lost 174k." Under Polish law (Art. 17 §1.11 of the PIT Act), a taxable event arises only when **crypto is exchanged for traditional currency or for goods**. As long as you hold a crypto position - regardless of how the market value moves - you don't report anything. Crypto-to-crypto swaps are also neutral (Art. 17 §1f). If you buy crypto for 100,000 PLN and don't sell for fiat for several years, that 100,000 PLN exists as a **documented cost of acquisition** (Art. 22 §14) and waits. In the year you sell crypto for fiat, that cost reduces the taxable base. If costs > revenue in a given year, the surplus **rolls forward indefinitely** (Art. 22 §16). That's where the _"buffer 174,895 PLN from 2024 → 162,948 PLN for 2026"_ comes from in my case. It's **not a loss** - it's expenditures on crypto purchases I haven't yet closed via sales to fiat. The buffer shrinks only when I actually sell crypto for PLN/EUR/USD. This is a heavily simplified description of the mechanism. Full understanding requires consultation with a tax advisor - this is not tax advice. ## Interpretive decisions: where the human comes back into play With 5,400+ transactions across multiple platforms, some categories of events have **ambiguous tax classification**. There are differing opinions from the Polish National Tax Information service (KIS), differing tax-advisor positions, and administrative-court rulings touching adjacent constructions. For each such category, the LLM laid out arguments for both sides, estimated risk exposure, and showed me the **numerical trade-off** - how much tax a conservative interpretation implies, how much a less restrictive one, and what the dispute risk looks like with the tax authority. **I made the decision**, not the agent. The agent left the rationale in `data/sources.md` in the repo - if an audit ever comes, I have a documented chain of reasoning. And one more key point worth repeating to anyone facing a first solo filing: **a PIT-38 amendment is possible up to 5 years back** (until 2030 for the 2025 return). Filing in good faith with documented decisions = a safe path. When you're not sure, a conservative interpretation in the first filing always works - you can later amend in your favour. In this section the LLM does not decide for me. It **maps options, documents arguments, waits for input**. That's exactly the division of labour I want. ## "You uploaded financial data to an LLM?" - a deliberate choice, not carelessness I know that for some readers, _"I uploaded financial data to an LLM"_ is already disqualifying. Three things to consider. **a) Claude Code (Anthropic API) does not use data for training models by default.** It's a different business model from consumer ChatGPT - and a qualitative difference. Current data-handling rules are always worth checking in the [Anthropic privacy policy](https://www.anthropic.com/privacy), but the baseline is this: the API is a production environment for developers and businesses, not a training-data harvester. **b) The repo is private and local.** There is no `git push` to GitHub. CSV files are in `.gitignore` - they don't even enter commit history. Raw financial data sits only on my disk. The LLM's context ends with the conversation. **c) A real benchmark.** The alternative was an accountant at a desk, an Excel file on a USB stick, or the tax-office portal in a browser - in every one of those scenarios my data passes through someone's hands, someone's memory, or someone's servers. **The choice is not between "safe" and "risky." It's between different kinds of trust.** I deliberately chose to trust Anthropic + a local workflow. Someone else will choose differently and that's fine. The discussion _"is an LLM safe for finances"_ loses meaning if we don't compare it to a concrete alternative. ## What I don't recommend - **I didn't automate the 172 PLN payment.** The transfer to the dedicated tax microaccount went manually. There's no point doing that through an LLM - one amount, one account number, 30 seconds in the banking app. - **I don't recommend this setup to anyone without programming comfort.** `/ingest`, directory structure, git, `.gitignore`, markdown editing - it requires familiarity with the tools. Without that, the time spent on setup eats all the gains. - **An LLM does not replace a tax advisor.** In ambiguous situations - and there are plenty when you have multi-source crypto + stocks + dividends - the final call has to belong to a human, ideally after a consultation with an advisor. The LLM provides arguments and maps risk. The advisor gives a recommendation tailored to your situation and takes responsibility for it. ## If you're reading this today - you still have ~30 hours If you haven't filed your PIT-38 yet and you have multi-source income, here's a minimum-viable plan for **2-3 hours** before the 30 April deadline: 1. Download every PIT-8C from your brokers (XTB, mBank, etc.) and reports from your crypto exchanges. 2. Open the official tax portal (Twój e-PIT) - section C (stocks + funds) is auto-filled for most people. 3. Sections E (crypto) and G (foreign dividends) have to be filled manually. This is where the portal won't help you. 4. **Check your PIT-38 from 2024.** See whether the relevant line carries unused crypto costs from prior years. That can be a buffer worth **a few to a few tens of thousands of PLN** that you might have forgotten about. 5. File via Trusted Profile (Profil Zaufany). Pay the microaccount by 30 April. 6. **Worst case: a PIT-38 amendment is possible until 2030.** Filing a roughly correct return on time is always better than no return + late-filing waiver paperwork. **After the deadline**, in three profiles: - **Simple PIT (one PIT-37 from employment)** - Twój e-PIT and that's it. Don't overcomplicate. - **2-3 sources (stocks + crypto)** - worth considering an evening or two to set up a workflow like this. - **5+ sources and a history of carried losses** - this is **your scenario**. The crypto cost-basis buffer can be worth 5-50k PLN in overlooked tax annually. The setup pays for itself. The value of an LLM grows not linearly but **in jumps**, when the repo's structures are legible to it. Without `CLAUDE.md` and project conventions, this project would have taken as long as working with an accountant. With them - two hours. That is exactly the ROI range that can't be sold via generic content marketing - because it requires the infrastructure to be there first.

Got multi-source taxes and thinking "I want this too"?

I help freelancers and technology consultants set up AI workflows for the kinds of work they used to delegate to experts. I'll show you what such a setup could look like for you.

Book a free consultation
## Useful Resources - [Twój e-PIT (Polish tax portal)](https://www.podatki.gov.pl/pit/twoj-e-pit/) - auto-fills section C for stocks and funds - [Polish PIT Act - ISAP](https://isap.sejm.gov.pl/) - Art. 17 §1.11, Art. 22 §§14, 16, Art. 11a (FX rates), Art. 30a (dividends) - [NBP - average exchange rates](https://nbp.pl/statystyka-i-sprawozdawczosc/kursy/) - D-1 rates for FX transactions - [Anthropic - Privacy & Data Usage](https://www.anthropic.com/privacy) - default API data handling - [Skills 2.0 - multi-agent system for company management](/en/blog/skills-2-0-multi-agent-system-company-management) - context for how `CLAUDE.md` conventions are built - [Spec-driven SEO on portfolio and Qamera AI](/en/blog/spec-driven-seo-portfolio-qamera-ai-case-study) - another case study, same workflow class ## FAQ
### Can I do PIT-38 with Claude Code if I'm not a programmer? Short answer: probably not worth it. The setup requires familiarity with git, the terminal, directory structures, markdown editing, and `.gitignore`. Without that, the time spent configuring eats all the gains. A safer path for non-technical users is an accountant or the official tax portal with manual entry of sections E and G.
### Does Anthropic use my financial data to train models? By default, no - Claude Code (Anthropic API) operates under a different business model than consumer ChatGPT. Data sent through the API is not used to train models without explicit customer consent. It's a different category from free consumer tools. Always worth checking the current rules in the [Anthropic privacy policy](https://www.anthropic.com/privacy).
### What is the "crypto cost-basis buffer" and why can it be worth tens of thousands of PLN? The buffer is your documented spending on crypto purchases that you haven't yet closed via sales to fiat currency. Under Polish PIT (Art. 22 §§14, 16) those costs wait in a buffer until the year you sell crypto for fiat - then they reduce the taxable base. Any surplus of costs over revenue in a given year rolls forward to subsequent years **with no time limit**. That's why checking last year's return can be worth several to several tens of thousands of PLN in overlooked buffer.
### What if I'm reading this after 30 April and haven't filed PIT-38? File as soon as possible with a "czynny żal" (active regret, Art. 16 of the Polish Fiscal Penal Code) - the penalty for non-filing grows with time, and czynny żal in many cases lets you avoid the fine. A PIT-38 amendment is possible up to 5 years back, so it's better to file a roughly correct return late than not at all. Afterwards, consult a tax advisor if your situation is complex.
### Does an LLM replace a tax advisor for multi-source PIT? No. An LLM is good at classifying transaction types, extracting data from PDFs, and mechanically applying NBP exchange-rate rules - but in cases of ambiguous legal classification the final call must belong to a human. Ideally - after a consultation with an advisor who takes responsibility for a recommendation tailored to your situation. The LLM maps options and risks. The advisor makes the decision and signs their name to it.
--- # Why WordPress can't do this - spec-driven SEO on a portfolio and Qamera AI Source: https://pawel.lipowczan.pl/en/blog/spec-driven-seo-portfolio-qamera-ai-case-study Published: 2026-04-26 ## Two projects, two stacks, one workflow loop In two weeks I optimized SEO on two radically different projects. **The portfolio** ([pawel.lipowczan.pl](https://pawel.lipowczan.pl)) - a Vite 7 + React 19 SPA with prerendering. The audit found 10 findings; I fixed five of them in a single afternoon. Securityheaders.com went from **C to A**, Rich Results Test from **5 warnings to 0**, the sitemap got per-URL `lastmod` instead of one build timestamp for all 73 URLs. **Qamera AI** ([qamera.ai](https://qamera.ai)) - my AI product photography SaaS, Next.js 16 App Router + Turborepo + Vercel + Supabase + i18n EN/PL/UK. The audit returned a health score of **56/100**. In five working days I closed **nine spec-driven changes** (seven planned plus two discovered along the way), which resolved every "Critical" finding. CLS on `/marketplace/styles` dropped from **0.467 to 0.016** - a 27× improvement. Hreflang coverage grew from "docs only" to **20 static marketing paths plus docs**. The homepage got three JSON-LD blocks (Organization + WebSite + SoftwareApplication); pricing got three more (Product × 2 + FAQPage). The central thesis of this article is simple and inconvenient for some readers: **full control over SEO and GEO is only possible with a code-based stack**. WordPress, Webflow, and Wix give you plugins - they don't give you a `Content-Security-Policy` header reporting to Sentry, sitemap-level `xhtml:link`, `requestIdleCallback` in ``, or `llms.txt` generated with your own build-time logic. The second multiplier is a good **AI workflow** - brainstorm → spec → execute → review → test. A code-based stack without a process = two weeks of manual work. An AI workflow on a closed platform = you hit the plugin ceiling. Both together = hours. I'll show you the process across both projects. You'll see what transfers 1:1 and what requires different decisions per stack. ## Why "platform vs code" is no longer a debate about hosting cost Five years ago, choosing WordPress over your own code was pragmatic. WordPress gave you themes, plugins, an ecosystem, and an admin panel for non-technical clients. Vercel with your own framework was overkill for 80% of projects. In 2026 those proportions shifted. **SEO has shifted toward GEO** (Generative Engine Optimization) - ChatGPT web search, Perplexity, Claude Search, and Gemini Deep Research read your pages, but differently from Googlebot. They respect the [llmstxt.org](https://llmstxt.org/) spec, weight `author.name` and `datePublished` in JSON-LD when picking sources, and prefer a "factual-definition opener" in the first 150 words. Security headers have become a trust signal. Core Web Vitals are validated with field data from CrUX, not lab scores from Lighthouse. Schema enrichment translates into rich results in SERP. What WordPress / Webflow give you in 2026: an SEO plugin (Yoast, Rank Math), basic schema for Article and Product, an automatically generated sitemap, redirects, `meta description`. That's about **80% of needs** for a typical company site. What they don't give you (or only with a serious fight): ```text - Content-Security-Policy Report-Only with reporting to Sentry - Permissions-Policy per-page (geolocation, camera, microphone, payment) - requestIdleCallback for third-party scripts instead of async=true - xhtml:link in the sitemap (not just hreflang in head) - llms.txt / llms-full.txt with your own generation logic - BlogPosting with mainEntityOfPage + publisher (raster logo) + ISO 8601 datetime - Per-bot rules in robots.txt (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) - Sitemap lastmod per page-type (post: frontmatter.modified, listing: max, legal: git mtime) ``` This is the top 20% of control. And it's exactly the range where today you win positions - in classic SERP and in LLM answers. Plugins close the top 80%. The top 20% requires editing HTTP headers, the HTML structure in ``, your own artifact builder. You don't do this in an admin panel - because the admin panel deliberately doesn't expose those layers. ## Toolchain - five tools, one loop I use the same five tools on every SEO project. Each on its own is ordinary. Together they form a loop that compresses time by 80-90%. | Tool | Role | |---|---| | **`claude-seo`** plugin (20+ sub-skills) | Audit as the first command - technical, GEO, schema, performance, hreflang, AI discoverability | | **OPSX / OpenSpec** | Spec-driven workflow - `proposal.md` → `design.md` → `specs/` → `tasks.md` before code | | **Lighthouse MCP** | Lab CWV and LCP opportunities from inside the agent, no context-switching to PageSpeed | | **Rich Results Test + securityheaders.com + Sentry CSP Reports** | Verification at every step | | **Git worktrees** | Parallel work on independent changes (when the project allows it) | I described OPSX in more detail in a [separate article on a structured approach to AI workflows](/en/blog/opsx-workflow-structured-ai-work). `claude-seo` is an example of a **specialized skill** in the sense from the [Skills 2.0 post](/en/blog/skills-2-0-multi-agent-system-company-management) - domain-specific knowledge plus checklist plus pre-built sub-skills. The work loop looks the same on every project: ```text ┌─────────┐ ┌──────────┐ ┌────────┐ ┌───────┐ │ Audit │ → │ Proposal │ → │ Design │ → │ Specs │ └─────────┘ └──────────┘ └────────┘ └───────┘ │ ┌─────────┐ ┌────────┐ ┌──────────┐ ┌───▼───┐ │ Archive │ ← │ Verify │ ← │ Implement│ ← │ Tasks │ └─────────┘ └────────┘ └──────────┘ └───────┘ ``` A typical command sequence: ```bash # Phase 1: audit /claude-seo:seo # Phase 2: change proposal /opsx:new "seo-improvements" /opsx:ff # fast-forward - all planning artifacts at once # Phase 3: implementation /opsx:apply # Phase 4: verification /opsx:verify # manually: securityheaders.com, Rich Results Test, Lighthouse MCP # Phase 5: archival /opsx:archive ``` Each of these tools works **because the substrate is code**. The audit can read any source file and the raw `` output. OPSX can edit `vite.config.js`, `next.config.ts`, `vercel.json`, `robots.txt` in one session. Verifiers get full output without a sandbox. This is not a coincidence - it's a consequence of architecture. ## The audit - what claude-seo finds on two radically different projects The first observation that surprised me: **claude-seo returns the same set of finding categories regardless of stack**. Only the way to fix them differs. | Project | Stack | Findings | Initial state | |---|---|---|---| | Portfolio | Vite 7 + React 19 + Vercel | 10 (4 perf, 2 schema, 2 security, 1 sitemap, 1 GEO) | securityheaders C, Rich Results 5 warnings | | Qamera AI | Next.js 16 + Turborepo + Vercel | 7 + 2 discovered along the way | health score 56/100 | Shared finding categories that transfer 1:1 between stacks: - **Missing or incomplete `llms.txt`** - biggest miss on GEO readiness in both projects - **Schema enrichment** - missing `publisher`, `dateModified`, `mainEntityOfPage`, ISO 8601 datetime - **Hreflang only at head-level**, no `xhtml:link` in the sitemap for clustering language variants - **Security headers** - deprecated `X-XSS-Protection`, no `Permissions-Policy`, no HSTS preload, no CSP - **AI bot allowlist is wildcard** - which for `Google-Extended` and `GPTBot` means "no signal", not "allow" Stack-specific findings that require different decisions: - **Portfolio:** `clickrank.ai` synchronous in `` blocks the parser before First Paint, sitemap with 73 URLs and a single `lastmod`, `articleBody: post.excerpt` semantically wrong in `BlogPosting` - **Qamera:** CLS 0.467 on `/marketplace/styles` from a client-side fetch from Airtable without reserved card dimensions, hardcoded EN strings in `root-metadata.ts` (PL/UK users got English OG on every marketing page), Merchant Listings false-positive on pricing The shared finding set is the **first proof that the process is transferable**. I know plugin-based SEO scanners exist, and each of them on the two projects would have returned fundamentally different reports because each is bound to a specific platform. An agent audit on a neutral substrate - code - returns a universal picture. ## From audit to change proposal - when to bundle, when to split Two projects, two different change-packaging strategies. **The portfolio** got one change `seo-improvements` with five pillars in a single PR ([#2](https://github.com/plipowczan/portfolio/pull/2)). Single-maintainer, no risk of file conflicts, easier review of the whole - because the changes are logically related thematically. **Qamera** got nine separate changes, eight PRs (#75/76/77/82/92/93/94/96), worked in parallel on git worktrees. Multi-developer, monorepo, disjoint file sets. Worktrees let each change have its own `node_modules` and its own dev server port - zero state conflict. The decision criterion I use: | Factor | One-PR (portfolio) | Multi-PR (Qamera) | |---|---|---| | Number of maintainers | 1 | 2+ | | File conflict risk | low | high | | Review cycle | self-review | peer review | | Time spread | one afternoon | 5 working days | | Rollback granularity | all or nothing | per-feature | | Dev environment isolation | not needed | worktree + separate node_modules | Common to both: every change = OPSX `proposal.md` + `design.md` + `specs/` + `tasks.md` **before** code. This isn't bureaucracy. It's a **feedback loop for AI**: reviewing a spec costs minutes, reviewing 200 lines of generated code in the wrong place costs hours. Spec-driven gives you a veto point before you pay the cost of implementation. ```text ## Tasks - seo-improvements - [x] Move clickrank inline to requestIdleCallback (+ setTimeout fallback) - [x] Generate dedicated raster logo (600×60 PNG via sharp) - [x] Add publisher / dateModified / mainEntityOfPage to BlogPosting - [x] Drop articleBody: excerpt (semantically wrong) - [x] Build llms.txt / llms-full.txt generator (scripts/generate-llms-txt.js) - [x] Replace X-XSS-Protection with Permissions-Policy + HSTS preload - [x] Configure CSP Report-Only → Sentry Security Reports bucket - [x] Per-page-type lastmod (post: frontmatter.modified, listing: max, legal: git mtime) ``` Each task in `tasks.md` is one unit of work with known scope and known verification. After implementation, the checkbox is proof the task was done - not a declaration. ## What transfers 1:1 (and why this is an argument for code) Four things I implemented in identical patterns on portfolio and Qamera. Each would be hard or impossible on a closed platform. ### A. `llms.txt` as your own build-time artifact The [llmstxt.org](https://llmstxt.org/) spec has existed since 2024 (Answer.AI / Jeremy Howard). In 2026 ChatGPT web search, Perplexity, Claude Search, and Gemini Deep Research respect it. The file is a shortened content index for LLMs, with an optional `llms-full.txt` containing the full content for single-token ingest. On the portfolio I have `scripts/generate-llms-txt.js` running in `build:prerender`. It reads `src/content/blog/*.md` (PL + EN) via `gray-matter`, plus `src/data/projects.js`. It generates `public/llms.txt` (~16 KB index) and `public/llms-full.txt` (~800 KB full content with `\n\n---\n\n` separator). ```text # Pawel Lipowczan > Software architect and technology advisor... ## Blog (PL) - [Title](url): one-line description ... ## Blog (EN) - [Title](url): one-line description ... ## Contact - email: ... ``` In Qamera a sibling script in the `apps/web` workspace generates `llms.txt` from marketing pages, blog posts, and public docs. The logic differs (data sources, section structure), but the pattern - a build-time generator following the spec - is identical. **You won't do this in a panel:** a WordPress plugin can spit out a static `llms.txt`, but you won't integrate it with your CMS on your own terms (section order per language, fallbacks for missing `description`, pagination at 100+ articles). ### B. Schema enrichment - `articleBody: excerpt` is a semantic error My old `BlogPosting` had six fields. Rich Results Test showed five non-critical warnings. After enrichment - eleven fields, zero warnings. ```json { "@type": "BlogPosting", "headline": "...", "description": "first 300 characters of content or frontmatter.description", "author": { "@type": "Person", "name": "Pawel Lipowczan", "url": "https://pawel.lipowczan.pl" }, "datePublished": "2026-01-15T00:00:00Z", "dateModified": "2026-04-21T00:00:00Z", "image": "...", "url": "...", "mainEntityOfPage": { "@type": "WebPage", "@id": "..." }, "publisher": { "@type": "Organization", "name": "Pawel Lipowczan", "logo": { "@type": "ImageObject", "url": "https://pawel.lipowczan.pl/logo-schema.png" } } } ``` Three non-obvious details: `articleBody: post.excerpt` is **semantically wrong** (the spec requires the full body, not a summary) - I removed the field entirely. `publisher.logo` must be a **raster** (PNG 600×60), not SVG. ISO 8601 with `Z` or offset, not `2026-01-15` without a timezone. In Qamera the same set of changes hit `Article` on `/blog`, `Service` on `/offer/*`, and `Product` on `/pricing`. **You won't do this in a panel:** SEO plugins set the top six fields. `mainEntityOfPage`, `publisher.logo` as a separate raster, ISO datetime, `description` with a fallback to the first paragraph - that's manual work in a schema generator. ### C. Hreflang at the sitemap level, not just `` Next.js `Metadata.alternates.languages` is a head-level signal. Google prefers **sitemap-level `xhtml:link`** for clustering language variants. In Qamera we solved it with a shared helper `buildLanguageAlternates(pathname)` used from two places - `sitemap.ts` and every page's `generateMetadata`. ```xml https://qamera.ai/pricing ``` A drift-guard test in CI fails when someone adds a path to the sitemap but forgets to add `alternates` to `page.tsx`. It's a safeguard against silent regression - it's very easy to add a new landing and forget about its language variant. **You won't do this in a panel:** Yoast generates hreflang in head. Sitemap-level requires editing the sitemap generator plus a drift-guard - i.e., CI code. ### D. AI bot allowlist - named rules instead of wildcard `robots.txt` with separate blocks for `GPTBot`, `OAI-SearchBot`, `ClaudeBot`, `PerplexityBot`, `Google-Extended`, and `CCBot`. Wildcard = "no signal" - the bot interprets it conservatively. Named allow = "explicit yes" - the bot knows it can crawl and index for its pipeline. **You won't do this in a panel:** WordPress writes to `robots.txt` via a plugin, but per-bot rules require editing the physical file - i.e., filesystem access you don't have on typical shared hosting. ## What's stack-specific (and what each project taught me separately) ### Portfolio - `async=true` on an inline script is a myth This finding surprised me the most, because everyone gets it wrong. The `clickrank.ai` script in `` looked like this: ```html ``` The trap: `async=true` applies to **fetching** the script, but the inline code that creates it executes **synchronously during HTML parsing**. It adds a microtask to the event loop before the browser renders anything. Fix - `requestIdleCallback` plus a fallback for Safari 16.3 and older: ```html ``` Verification after deploy to prod: `performance.getEntriesByType('resource').filter(r => r.name.match(/clickrank/))` → `startTime: 101.6ms`. The browser reported idle after ~100ms and only then fired the callback. Lighthouse lab variance stayed large (post: prod 38 → preview 61 → second run 43) - **lab score ≠ field data**. Real verification is CrUX from Google Search Console after 2-4 weeks. ### Qamera - CLS 0.467 → 0.016 via SSR initial grid `/marketplace/styles` rendered style cards loaded client-side from Airtable. Without reserved dimensions, the grid layout shifted 4× over the failing threshold (0.467 vs target ≤ 0.1). Three options for the fix - SSR initial grid, reserved card dimensions, combined. We chose **SSR**: bonus for GEO (non-JS crawlers see the content) plus eliminating CLS at the source. Result: **CLS 0.016** (27× improvement), LCP 2.4s → 1.6s. Post-deploy twist: PageSpeed Insights showed LCP 14.4s (cold Vercel function), Lighthouse MCP in parallel: 1.6s (warm). **A single PSI metric is sampling.** Always re-run or verify locally. ## A bug the audit wasn't looking for - and why this argues for regular audits The portfolio's SEO audit surfaced a finding I didn't expect: hreflang alternates for the post `llm-knowledge-base-brain-karpathy` pointed to `/en/blog/`, which returns "Post not found". Root cause: the post was **PL-only** (no EN version), but its frontmatter had `alternateSlug: llm-knowledge-base-brain-karpathy` - pointing at itself. The result of an earlier iteration of the blog-article-writer skill that **auto-filled the field without validation**. The chain of events: 1. User on the PL post clicks the language switcher 2. `getAlternatePost(currentSlug)` returns... the same PL post 3. LanguageSwitcher builds `/en/blog/` and navigates 4. `BlogPostPage` filters `getPostsByLang("en")` → no match → "Post not found" 5. The sitemap inherits this bug as bad hreflang and propagates to Google A three-level fix: **data fix** (remove the field), **code defense** (`getAlternatePost` rejects self-reference and same-lang candidates), **process fix** (rule in `.claude/rules/data-storage/` plus updating the validator in the blog-article-writer skill). The meta-lesson is strong: an SEO audit triggers bugs **that weren't its target**. I would never have found this without claude-seo. That's an argument for regular audits even on a small project. The second meta-level - the bug was **introduced by an AI workflow** (the blog-article-writer skill), fixed by **another AI workflow** (audit + spec-driven fix + a rule in the skill). That's a self-correcting loop, provided there's a process. Without a process - the bug would have lived for weeks. A related thread on how an agent enforces its own standards I described in the [Karpathy LLM Wiki post](/en/blog/karpathy-llm-wiki-knowledge-base) and [Second Brain with Obsidian and Claude Code](/en/blog/second-brain-obsidian-claude-code-skills). ## Time compression is multiplicative, not additive By the numbers: - **Portfolio:** audit 15 min + 4h implementation + 30 min verification = **5h** for 5 changes - **Qamera:** **5 working days** for 9 changes (seven planned plus two discovered along the way) - **Second project = ~30% the time of the first** thanks to pattern transfer (llms.txt, schema enrichment, hreflang sitemap, AI bot allowlist) Multiplier: **code-based stack × good AI workflow = hours**. Each on its own isn't enough. A code-based stack without a process = two weeks of manual work with forums and Stack Overflow. An AI workflow on a closed platform = you hit the plugin ceiling in half an hour. Together they yield 80-90% compression. That's not additive - it's multiplication. The argument for why I increasingly choose code-based solutions I also described in the [vibe coding guide](/en/blog/vibe-coding-guide). Six takeaways from both projects: 1. **A code-based stack gives you the top 20% of control plugins don't** - and that's the range that wins positions today 2. **Spec-driven as a feedback loop for AI** - reviewing a spec costs minutes, reviewing 200 lines of code costs hours 3. **Transferable 1:1 between stacks:** llms.txt, schema enrichment, hreflang sitemap-level, AI bot allowlist 4. **Stack-specific:** every framework has its own performance traps and its own metadata API - that's where you save least 5. **The audit finds bugs outside its scope** - `alternateSlug === slug` was never on the list, I found it via claude-seo 6. **Second project = 30% the time of the first** - provided you document patterns If you stay on WordPress, this article doesn't change your life. If you're considering moving to your own stack, it's the argument you needed.

Need an SEO + GEO audit on your own stack?

I do the same on client projects - from audit through spec-driven changes to post-deploy verification. We'll discuss your stack and a realistic scope in 30 minutes.

Book a free consultation
## Useful Resources - [llmstxt.org](https://llmstxt.org/) - llms.txt spec - [securityheaders.com](https://securityheaders.com/) - security headers scanner - [Rich Results Test](https://search.google.com/test/rich-results) - Google's structured data validator - [PageSpeed Insights](https://pagespeed.web.dev/) - Core Web Vitals lab + field data - [MDN - requestIdleCallback](https://developer.mozilla.org/en-US/docs/Web/API/Window/requestIdleCallback) - [MDN - Content-Security-Policy](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Content-Security-Policy) - [Sentry Security Reports](https://docs.sentry.io/product/security-policy-reporting/) - CSP reports via Sentry - [Google Search Central - Article structured data](https://developers.google.com/search/docs/appearance/structured-data/article) - [Qamera AI](https://qamera.ai) - the project described in the case study ## FAQ
### Is every code-based site automatically better at SEO than WordPress? No - a code-based stack gives you **control**, not results. Without a process (audit → spec → execute → verify), you'll end up with a worse site than a well-configured WordPress with Yoast. The argument of this article is this: a code-based stack **lets you** optimize the top 20% (CSP, llms.txt, sitemap-level hreflang, schema enrichment) that platforms don't expose. Whether you use it depends on your workflow.
### What does spec-driven development mean in the context of SEO? Spec-driven means every change starts with artifacts: `proposal.md` (what and why), `design.md` (how), `specs/` (contracts), `tasks.md` (a list of steps) - **before** writing code. It works especially well in SEO because changes touch many layers (HTTP headers, HTML head, structured data, sitemap), and skipping the spec means AI generates 200 lines of code in the wrong place. I use the OpenSpec / OPSX workflow - details in a [separate article](/en/blog/opsx-workflow-structured-ai-work).
### Does llms.txt make sense in 2026 if my site isn't an AI tutorial? Yes, but the ROI is lower. `llms.txt` works strongest for content that LLMs cite (tutorials, documentation, case studies). For e-commerce or portfolio the impact is smaller but still positive - the cost is 100-200 lines of a Node script, the benefit is presence in ChatGPT, Perplexity, and Claude Search grounding. The `llms-full.txt` file gets heavy at 100+ articles - at that scale, paginate or use `top-articles-only`.
### How do I choose between one big PR and many small ones for SEO changes? Single-PR makes sense with a single maintainer and thematically coherent changes (like the portfolio: 5 SEO pillars in one PR, 4h of work). Multi-PR is necessary with multiple maintainers, monorepo, and parallel work (like Qamera: 9 changes, 8 PRs, 5 days). The criterion: do the changes touch the same files (conflict = split) and is reviewing the whole realistic in one pass (>500 lines diff = split).
### Does an AI workflow replace code review on SEO changes? No - it complements it. On Qamera, Copilot review on a PR caught three valid issues (placeholder Sentry DSN, missing preview env var, unused import) that the spec-driven workflow didn't catch. An AI workflow speeds up generating code that conforms to the spec, but a **second pair of eyes** (human or AI reviewer) still catches the difference between "the code does what the spec says" and "the code does what the spec says, in a way that's safe for production".
--- # How Karpathy's LLM Wiki Helped Me Organize My Knowledge Base Source: https://pawel.lipowczan.pl/en/blog/karpathy-llm-wiki-knowledge-base Published: 2026-04-12 ## Karpathy described a framework. I had a living system to organize In April 2026, Andrej Karpathy published [a thread about "LLM Wiki"](https://x.com/karpathy/status/2039805659525644595) on X - a concept where the LLM builds and maintains a persistent wiki from your sources. Instead of classic RAG, which retrieves fragments on demand, the LLM actively manages the knowledge base: it creates notes, updates cross-references, flags contradictions. I read the thread and had déjà vu. Not because I'd built the same thing - but because, since 2022, I'd been organically evolving in that direction. My Obsidian vault started as a classic pile of hand-written notes - a few dozen Markdown files, manually linked, growing without a clear structure. Over time, **Claude Code** took over more and more of the work: first simple formatting, then indexing, eventually full management of structure and standards. At some point I realized the agent was doing more maintenance than I was. Karpathy's repository - [LLM Wiki gist](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) - wasn't the discovery of a new world for me. It was a **catalyst for tidying up**. Karpathy gave names to the layers that had existed chaotically in my setup. He helped me formalize the separation between raw sources and the processed wiki. To write down safety rules I'd only been keeping in my head. He described a navigation protocol that worked intuitively in my system but had never been documented. > _"Boring part of maintaining a knowledge base isn't reading or thinking - it's bookkeeping."_ - Karpathy This is the story of an evolution from manual PKM to an agent-powered knowledge system. With Karpathy as the catalyst who helped me organize what I'd been building for years. And with one central concept that changed everything: **progressive disclosure** - organizing notes so an AI agent can find them without cluttering its context window. In this article I'll walk through that evolution: from the chaos of hand-written notes, through gradually handing maintenance to the agent, to a system with 185 notes, three navigation indexes, and four workflows. I'll show the actual architecture, the actual processes, the actual numbers - not theory. ## The problem with classic PKM Anyone who has tried to maintain a knowledge base for more than a few months knows the pattern: enthusiasm at the start, growing chaos in the middle, abandonment at the end. Why classic approaches don't scale: - **Maintenance burden grows faster than value.** Every new note is a potential link to dozens of existing ones. At 50 notes I can still handle it. At 185 - no chance. The cost of keeping order grows exponentially; the value of each added note grows linearly. - **Cross-references are always incomplete.** I write a note about context engineering and forget that three months ago I saved something related in a completely different folder. The link never gets created. And without that link, the knowledge is fragmented - as if it didn't exist. - **Wikis die from maintenance, not from lack of value.** Nobody abandons a knowledge base because it's worthless. They abandon it because upkeep becomes sickening. I've seen it more than once - in companies, in open source, in my own projects. - **Zettelkasten, BASB, Digital Garden - beautiful philosophies.** Tiago Forte's Building a Second Brain, Luhmann's Zettelkasten, Digital Garden - each of these methodologies makes sense on paper. But they all assume YOU are the maintenance bottleneck. And they're right - which is why systems die. Not because people are lazy. Because bookkeeping isn't work you have time for alongside writing code, running a business, and reading new things. Karpathy put it precisely: the boring part of maintaining a knowledge base isn't reading or thinking. It's bookkeeping - updating cross-references, keeping summaries current, noting contradictions. **LLMs don't get bored.** And that's the key observation. It's not that AI is smarter at organizing knowledge. It's that AI has no problem with the repetitive, boring tasks humans put off until "someday." Over three years with Obsidian, I went through every one of those stages. I started enthusiastic, then came the duplicates and orphan notes, and eventually I hit the point where finding anything took longer than writing it from scratch. My vault survived because at some point I stopped fighting maintenance manually - and handed it to the agent. ## The architecture - what it looks like after the cleanup ### Tech stack The system rests on five pieces: - **[Obsidian](https://obsidian.md/)** - editor and frontend (local, Git-based) - **[Quartz 4](https://quartz.jzhao.xyz/)** - Static Site Generator → GitHub Pages - **GitHub Actions** - automatic deploy on push to the `v4` branch - **Claude Code** - agent maintaining the vault - **Public URL:** [brain.lipowczan.pl](https://brain.lipowczan.pl) Obsidian is the interface for me - this is where I read, browse graph view, jot ad hoc notes. Quartz 4 is the interface for the world - a static site on GitHub Pages, available at [brain.lipowczan.pl](https://brain.lipowczan.pl). Claude Code is the engine that keeps order between the two. The whole thing is a Git repo - every change tracked, every ingest a commit, full revertable history. ### Three layers - Karpathy's framework that helped me organize These layers had existed organically in my setup for a long time. I had an inbox for raw material, processed content, and some form of configuration. But Karpathy gave them names and a formal structure - and that helped me sharpen them. | Layer | Karpathy (LLM Wiki) | My implementation | | ----------- | ----------------------- | -------------------------------------------------- | | Raw sources | Immutable drop zone | `/content/_raw/inbox/` - drop zone, out of build | | Wiki | LLM-generated .md files | `/content//` - built and published | | Schema | Config document | `CLAUDE.md` - 300+ lines of agent configuration | Separating raw sources from the wiki is the critical change I formalized after reading Karpathy. Before, raw material and processed notes mixed in one directory - sometimes the agent modified the source instead of creating a new note. Now the inbox is an immutable drop zone - the agent processes, but the original stays untouched in `_raw/processed/`. I can always go back to the source and compare it with what the agent made of it. It's a simple safety mechanism, but it gives peace of mind - I know nothing is lost. ### Progressive disclosure - why notes are organized FOR the agent This is the central concept of the whole system. Notes aren't organized to look pretty in Obsidian's graph view. They're organized so the **AI agent can quickly find them WITHOUT cluttering its context window** with unnecessary information. Three navigation indexes - from general to specific: - **`vault-map.md`** (~80 lines) - a bird's-eye view of the entire vault. The agent ALWAYS reads this first. Enough to understand the structure and decide where to look next. - **`catalog.md`** (~650 lines) - one line per note with title, category, and a short description. The agent reads this when it needs to find a specific note. - **`graph.md`** - wikilink graph (outgoing + incoming edges). The agent reads this when it needs context on the connections between notes. ```text _indexes/ ├── vault-map.md ← always first (~80 lines) ├── catalog.md ← 1 line / note (~650 lines) └── graph.md ← wikilink edges ``` The agent navigates like a human with a table of contents - it doesn't grep 185 files. It reads vault-map (80 lines), decides which topic is relevant, reaches into catalog for specific notes, and if it needs context on connections - it opens graph. That keeps the context window clean, and answers get sharper. Compare that with the naive approach: "dump all files into the context window and ask." At 185 notes of ~500 words each, that's ~92,500 tokens of notes alone. Most models either can't fit that, or they "get lost" halfway through - they start answering based on whichever fragments happen to land near the query in embedding space, ignoring the rest. Progressive disclosure solves this elegantly: the agent reads only as much as it needs, in order from general to specific. That's **progressive disclosure** in practice. Minimizing context window usage isn't optimization - it's the foundation without which the agent drowns in noise. Without navigation indexes you have two options: either hand the agent the whole vault (too much context) or point it at specific files yourself (back to manual work). The indexes eliminate both problems. ## CLAUDE.md - "schema" as the heart of the system Karpathy calls this the **"schema document"** - a file that defines how the agent should treat the knowledge base. In my case it's `CLAUDE.md` - 300+ lines of configuration injected into Claude Code's system prompt at every session. What my CLAUDE.md contains - and why each section exists: - **Directory structure** and the purpose of each folder - so the agent knows where to create notes and what not to touch - **Navigation protocol** - progressive disclosure, the order for reading indexes, when to reach for full files - so the agent doesn't pollute its own context window - **Writing style guidelines** - note writing style, EN/PL mix, emoji in headers, formatting - so all notes look consistent regardless of session - **Frontmatter schema** - YAML metadata for each note (required fields, types, validation) - so the agent doesn't create notes with incomplete metadata - **Workflow definitions** - INGEST, COMPILE, LINT, Q&A, ENHANCE with concrete steps - so every workflow is repeatable and deterministic - **Safety rules** - what the agent should never modify or delete (e.g., indexes only updated, never rebuilt from scratch; don't change frontmatter on existing notes without confirmation) ```yaml # Excerpt from the frontmatter schema in CLAUDE.md required_fields: - title # Note title - category # Topical category (AI, BUSINESS, CODE...) - tags # Tag list - summary # 2-3 sentence summary - created # Creation date (YYYY-MM-DD) - updated # Last update date - source # Where the knowledge came from (URL, book, experience) ``` **Key insight:** CLAUDE.md isn't generated by an LLM. It's written by hand and evolves with experience. An ETH Zurich study (2026) showed that LLM-generated agentfiles _degraded_ agent performance while costing **20%+ more**. The reason is simple - LLMs generate verbose, redundant instructions that clutter the context window. Hand-written instructions are precise and universally applicable. It's counterintuitive. It seems like if an LLM writes better prose than most people, it should write its own configuration. In practice - the LLM generates "just in case" instructions, repeats itself, adds edge cases that never occur. The result: 800 lines instead of 300, context window polluted, agent slower and less precise. > _"Key insight: most agent failures aren't a model problem - they're a configuration problem."_ Less is more. Every line in CLAUDE.md must justify itself. I start simple, add only what I actually need, remove what doesn't work. It's a living document - not a spec written once and forgotten. My CLAUDE.md has been through a dozen-plus iterations - some rules I added after the agent deleted important metadata, others after it ignored existing connections. Every rule is a lesson from a concrete failure. ## Workflows - what the agent can do ### INGEST - from file to knowledge in 5 minutes This is the most common workflow - and the one that best shows the system's value. Processing a single article into a note used to take me 20-30 minutes. Ten sources - a whole evening. With the agent? 5 minutes plus review. What INGEST looks like step by step: 1. **Drop** - I drop a file into `_raw/inbox/` - an article, PDF, transcript, Web Clipper note, anything in a text format 2. **Trigger** - I type `ingest` in Claude Code 3. **Navigate** - the agent reads `vault-map.md`, checks `catalog.md` for overlap with existing notes. If the topic already exists - it proposes an update instead of a duplicate 4. **Create** - creates a note from the template, fills in frontmatter (category, tags, summary, source, date) 5. **Link** - adds wikilinks to related notes - and updates those notes to link back (bidirectional linking) 6. **Archive** - moves the source to `_raw/processed/YYYY-MM-DD_name` - the original stays, just out of the inbox 7. **Index** - updates all three indexes (vault-map, catalog, graph) A single source can touch **10-15 files** in one pass. The note itself plus updates to related notes plus three indexes. By hand - 2 hours if you care about cross-references. With the agent - 5 minutes plus my review, which checks whether the agent understood the material correctly and placed it in the right context. The review is critical. I don't accept INGEST output blindly. I check: is the category correct? Do the tags make sense? Do cross-references point to genuinely related notes, not loose associations? Does the summary capture the essence of the source? That takes 2-3 minutes, but it gives confidence that the vault keeps its quality. The agent does 95% of the work - I take care of the crucial 5% that requires human judgment. ### COMPILE - synthesize from many sources The agent combines multiple notes into a compiled article. Example: my note on LLM Knowledge Bases is a compiled article pulled from Karpathy's thread, my notes on RAG, and context from the vault. The agent reads the sources, identifies common threads, synthesizes, and creates a `compiled-note` with full citations. What that looks like in practice: I say "compile everything I have about context engineering into a single note." The agent reads the catalog, identifies 7 related notes, reads them, extracts the key concepts, and builds a coherent article with sections and citations. Result: one comprehensive note instead of 7 scattered fragments, with clear attribution to sources. This is a workflow I'd never do regularly by hand - because it requires reading and cross-referencing several documents at once. Seven notes of 500 words each is 3,500 words to read, understand, and synthesize. The agent does it in minutes - and crucially, it doesn't skip notes I'd miss because I forgot they existed. ### Q&A - answers with citations in 30 seconds "What do my notes say about context engineering?" - the agent reads the indexes, identifies candidates, reads the notes, synthesizes an answer with citations. Good answers can flow back into the wiki as new compiled-notes. This turns the knowledge base from a passive archive into an **active thinking tool**. Instead of digging through folders, I ask a question and get a synthesis from my own notes - with links to sources. The crucial part: the agent answers based on MY knowledge, not the general internet. If three months ago I captured an important insight from a conference - Q&A will find it and bring it back up, even if I forgot I'd written it down. Another example: "Compare what my notes say about RAG vs what Karpathy describes as LLM Wiki." The agent reads both notes, identifies points of agreement and disagreement, and presents the comparison. That's the kind of analysis which, by hand, would require opening two documents and comparing paragraph by paragraph. ### LINT - health check A periodic health check of the vault: broken wikilinks, orphan notes (notes with no connections), missing summaries, TODO markers, incomplete frontmatter. The agent generates a report into `_outputs/reports/` and proposes fixes. A typical LINT report after a month: - 4 broken wikilinks (notes moved/renamed without updating references) - 2 orphan notes (added but never linked to the rest of the vault) - 7 missing summaries (notes with an empty summary field in frontmatter) - 3 notes with TODO markers (unfinished processing) LINT is a workflow you don't think about until you see a report of 23 broken wikilinks after a month of adding notes. It's hygiene - like code linting. Not sexy, but without it a vault degrades in silence. The agent does it for me - regularly, without forgetting. ## From hand-written notes to 185 files maintained by the agent The evolution of this system wasn't planned. I didn't sit down in 2022 with an architecture in mind. It was an organic process where each phase solved a concrete problem from the previous one. **Phase 1 (2022-2024): Hand-written notes in Obsidian.** Classic setup - topical folders, wikilinks, tags. I captured notes from books, conferences, courses. Worked up to ~50 notes. After that, growing chaos: incomplete cross-references, forgotten content, duplicates. Searching for specific information started taking longer than just Googling it again from scratch. The typical PKM trajectory - enthusiasm, plateau, frustration. **Phase 2 (2024-2025): The agent starts helping.** When I started working intensively with Claude Code, I naturally began using it for simple tasks in the vault. First formatting and filling in frontmatter - boring work, perfect for an agent. Then linking new notes to existing ones - the agent is better at this than I am, because it reads the full catalog every time. Finally - full source processing from scratch: drop a file, say `ingest`, the agent handles the rest. This is the phase when the first CLAUDE.md was born - still simple, ~80 lines, but already defining the structure and core rules. **Phase 3 (2026): Karpathy LLM Wiki → formalization.** Karpathy's thread gave me a framework to name what I already had. Three layers - raw sources, wiki, schema - suddenly had official names. The navigation protocol stopped being "the way I do it" and became a documented procedure. Safety rules I'd kept in my head landed in CLAUDE.md. I formalized the processes, filled in missing pieces, tidied up the indexes. CLAUDE.md grew from 80 to 300+ lines. Where I stand today: **185 notes**, 13 topical categories. Breakdown: - `AI/` - 19 notes (Claude Code, harness engineering, context engineering, skills) - `LIFE/` - 40 notes (books, knowledge, tools) - `BUSINESS/` - 26 notes - `CODE/` - 23 notes - `PROJECTS/` - 12 notes (Qamera AI, Brain, Agentic Systems) Workflow comparison - before and now: - **Before:** read an article → maybe save the link in bookmarks → forget where and in what context → search again when I need it → lose 20 minutes reconstructing context - **Now:** click Obsidian Web Clipper → the article lands in `_raw/inbox/` → I type `ingest` in Claude Code → the agent weaves the note into the knowledge network with cross-references → when I ask: Q&A in 30 seconds with citations and links to sources > _"Wiki is a persistent, compounding artifact - cross-references ready, contradictions flagged."_ What compounds: the agent guards structure and standards, every note lands in the right place with the right links. Knowledge accumulates, it doesn't scatter. The more notes, the richer the network of connections - and the more valuable each next note, because the agent has more context to link against. It's a reversal of the traditional PKM problem, where more notes = more chaos. Here, more notes = a richer graph. ## Notes organized for the agent, not for aesthetics This is the key mindset shift that changed my approach to PKM. You don't organize notes so YOU can search them more easily. You organize them so the **AGENT can find them more easily**. The agent is the primary consumer of your knowledge base. You are the curator - responsible for sourcing (what you drop into the inbox), exploration (what questions you ask), and decisions (what to compile, what to archive). The agent handles the bookkeeping: creating notes, updating indexes, watching cross-references, flagging contradictions. That division lets you focus on thinking, not on administration. **Progressive disclosure** minimizes context window usage. The agent doesn't have to read 185 files to answer a question. It reads 80 lines of vault-map, identifies the relevant area, reads a few specific notes. Sharper answers, lower cost, faster execution. To see it in numbers: if the agent read the whole vault on every query, that's ~185 files × ~500 words = ~92,500 words injected into the context window. With progressive disclosure: vault-map (80 lines) + 3-5 relevant notes = ~3,000 words. **30x less context, but sharper answers** - because the agent knows what it's reading and why. The parallel with agentic coding is direct: | | Agentic Coding | LLM Knowledge Base | | ----- | ------------------------------------------------- | ------------------------------------ | | You | Environment architect (specs, context, guardrails) | Curator (sourcing, questions, decisions) | | Agent | Implements code | Bookkeeper, writer, cross-referencer | The agent doesn't replace thinking - it replaces bookkeeping. That's not a subtle difference. The decision about WHAT to ingest, what questions to ask, what to compile - that's still your job. The agent is a machine for maintaining order, not a machine for thinking on your behalf. When you try to hand the agent decisions - you get generic, "safe" answers. When you hand over maintenance and focus on curation yourself - you get a system that compounds with every new note. Every source enriches the existing knowledge network. Cross-references form automatically. Contradictions between notes get flagged. The summary is always up to date. > _"Shift: from 'developer who writes code' to 'architect who designs systems for agents to write code.'"_ The same shift applies to knowledge: from "person who maintains notes" to "curator who designs systems for agents to maintain knowledge." You decide what's worth saving. The agent takes care of making sure what you saved stays accessible, linked, and current. ## Takeaways Five conclusions after three years of building this system: 1. **Methodology > tool.** Zettelkasten, Building a Second Brain, LLM Wiki - these are philosophies, not apps. Obsidian, Notion, Logseq - these are tools. What matters is understanding _why_ you organize knowledge a certain way, not _what in_. You can build the same system in Notion with an agent, or in plain VS Code with Markdown files. The tool is interchangeable - the rules of navigation, indexing, and progressive disclosure work everywhere. 2. **CLAUDE.md is a contract with the agent.** It evolves from experience - you start simple, add based on what the agent does wrong, remove what doesn't work. After a month you have a solid configuration. After a year - a mature system. 3. **Progressive disclosure works.** Three index levels (vault-map → catalog → graph) is not over-engineering. It's the only thing that lets the agent operate sensibly over 185+ notes without grepping the whole directory and polluting its context window. 4. **The agent doesn't replace thinking.** It replaces bookkeeping. That's a fundamental difference. When you try to hand the agent _decisions_ - you'll get generic answers. When you hand over _maintenance_ - you'll get a system that compounds. 5. **A Git repo as the foundation.** Version history, branching, collaboration - for free. Quartz 4 as an SSG gives you a public digital garden on GitHub Pages. The entire vault is a Git repo - every change tracked, revertable, diffable. If the agent does something bad (and it will, especially early on) - `git diff` shows what it changed, `git revert` undoes it. A safety net without which I wouldn't hand over control of 185 files. If you want to start from zero and don't have a vault with an agent yet, read [Second Brain with Obsidian and Claude Code](/en/blog/second-brain-obsidian-claude-code-skills) - I describe there how to start from installing Obsidian through your first Skills. This article shows where you can get to after a few years of iteration - from a simple vault with a dozen notes to a system with 185 files, three layers, and four workflows, maintained by an AI agent. You don't have to build all of this at once. Start with a CLAUDE.md of 50 lines, one workflow (INGEST), and one index (vault-map). The rest will come organically - the way it came for me.

Want to build a similar knowledge management system?

I'll help you design a knowledge base architecture with an AI agent - from the CLAUDE.md structure through navigation indexes to workflows tailored to your needs.

Book a free consultation
## Resources - [Karpathy's X thread](https://x.com/karpathy/status/2039805659525644595) - the original post about LLM Wiki (April 2026) - [LLM Wiki gist on GitHub](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) - configuration and documentation - [brain.lipowczan.pl](https://brain.lipowczan.pl) - my public vault (Quartz 4 + GitHub Pages) - [Quartz 4](https://quartz.jzhao.xyz/) - Static Site Generator for Obsidian vaults - [Obsidian](https://obsidian.md/) - Markdown note editor - Related article: [Second Brain with Obsidian and Claude Code Skills](/en/blog/second-brain-obsidian-claude-code-skills) - how to start from zero - [second-brain-template](https://github.com/plipowczan/second-brain-template) - a ready template to start your own second brain ## FAQ
### How does LLM Wiki differ from traditional RAG over documents? RAG (Retrieval-Augmented Generation) fetches text fragments on demand - you get an answer based on the nearest embeddings, but without broader context. LLM Wiki is a persistent knowledge base in which the agent actively builds and maintains notes: it creates cross-references, synthesizes knowledge from multiple sources, and flags contradictions. It doesn't answer a single question - it maintains a complete, evolving knowledge system.
### Do I need Obsidian to build a similar AI-powered knowledge management system? No - Obsidian is a convenient frontend with graph view and plugins, but it's not required. Any text editor plus a Git repo with `.md` files is enough. The key pieces are CLAUDE.md (the schema document defining how the agent should treat the vault) and the navigation indexes (vault-map, catalog, graph). You can build the same system in VS Code, Cursor, or even straight from the terminal.
### How long does it take to write a CLAUDE.md for a knowledge base, and how do I start? The first version takes 1-2 hours - you define the directory structure, a basic frontmatter schema, and one workflow (INGEST is a good choice). But CLAUDE.md is a living document that evolves with experience. You start with simple rules, watch what the agent does wrong, add corrections. After a month of regular use you'll have a solid configuration. The key rule: write it by hand, don't generate it with an LLM.
### What is progressive disclosure and why is it critical for an AI-agent knowledge base? Progressive disclosure is the principle of organizing notes in layers - from a general table of contents (vault-map, ~80 lines) through a catalog (catalog, ~650 lines) down to specific files. The agent reads only as much as it needs, without cluttering the context window with unnecessary information. That's the key difference between a knowledge base that "works" and one where the agent gets lost among 185 files and returns generic answers.
### Does an LLM Wiki system require a technical background to set up? The basic setup requires familiarity with Git, the terminal, and Markdown files - you don't need to code. The hardest part is writing a good CLAUDE.md that precisely defines how the agent should treat your vault. This article describes a mature system after 3 years of iteration, but you can start small - a single folder, a simple frontmatter schema, and the INGEST workflow. The article [Second Brain with Obsidian and Claude Code](/en/blog/second-brain-obsidian-claude-code-skills) gives you a solid starting point.
--- # 5 GitHub Repositories That Will Transform Your Work with Claude Code Source: https://pawel.lipowczan.pl/en/blog/5-github-repos-claude-code Published: 2026-03-31 # 5 GitHub Repositories That Will Transform Your Work with Claude Code Using Claude Code without the skills ecosystem is like using a smartphone without apps. Sure, it works, but you're leaving most of its potential on the table. I went through this myself - I'd open the terminal, type prompts, get code back. Sometimes good, sometimes generic. No system, no repeatability. Then I started exploring GitHub and discovered that a powerful ecosystem has grown around Claude Code. Skills, frameworks, integrations - hundreds of repositories, each promising a revolution. The problem? Most of it is noise. It's hard to separate tools that actually make a difference from those that look great in the README but don't hold up in day-to-day work. In this article, I've gathered **5 GitHub repositories** that truly changed how I work with Claude Code. I've tested each one in a production workflow - I don't recommend things I don't use myself. ## 1. UI/UX Pro Max - No More Generic AI Slop You know the problem. You ask Claude Code to build a page and you get... exactly the same thing as everyone else. The same layout with a hero section, the same rounded cards, the same gradients. **Generic AI slop** - that cookie-cutter look that immediately reveals the page was generated by AI. **UI/UX Pro Max** is a skill that solves this problem at its root. Instead of one universal approach to design, it offers **intelligent design system generation** tailored to what you're actually building. ### How It Works The skill analyzes the project type - portfolio, SaaS, e-commerce, landing page - and selects the appropriate design system. Different colors, different proportions, different components. This isn't a random pick. Each system has its own logic: a portfolio emphasizes personal branding, SaaS focuses on conversion, e-commerce on product presentation. In practice, this means two different projects generated with the same skill look **completely different**. And that's exactly the point - individuality instead of templates. ### My Experience I use UI/UX Pro Max in combination with **Tailwind CSS** and **React**. When building components for clients, the skill generates a cohesive design system that I then customize. It saves me time at the prototyping stage - instead of starting from scratch or fighting with generic output, I have a solid foundation tailored to the project context. If you're building anything with a frontend and want it to look professional without hiring a designer - this is your starting point. **Repository:** [UI/UX Pro Max](https://github.com/nextlevelbuilder/ui-ux-pro-max-skill) ## 2. OpenSpec - Structured Development Instead of Chaos **OpenSpec (OPSX)** is a framework for spec-driven development that truly changed how I work with Claude Code. I use it daily. ### The Problem It Solves Anyone who's worked with Claude Code for more than a week knows this scenario: you start a session, build a feature, the context grows, and the agent starts "forgetting" earlier decisions. This is **context window rot** - the degradation of response quality as the conversation gets longer. OpenSpec solves this through **spec-driven development**. Instead of chaotic sessions where you tell the agent what to do step by step, you create structured artifacts: a change specification, an implementation plan, delta specs. The agent knows what it's building, why, and how - before it writes the first line of code. ### What the Workflow Looks Like A typical session with OpenSpec starts with **explore** - and this is the crucial step most people skip: ```text 0. /opsx:explore -> brainstorming with the agent 1. /opsx:new "add dark mode to the blog" -> creates a change with artifacts 2. /opsx:ff -> fast-forward through all artifacts 3. /opsx:apply -> implement tasks from the plan 4. /opsx:verify -> verify against the spec 5. /opsx:archive -> archive the completed change ``` The explore phase is when the agent bounces ideas with you, asks for details, proposes approaches. This is where the decision happens - whether you even need a full specification, or if a simple proposal with a task list is enough. A simple change doesn't need an elaborate spec. A complex feature - absolutely. From my experience: **the more time you spend on preparing a good specification, the fewer iterations you'll need during the actual code implementation**. It pays off many times over. Each step produces a concrete artifact - a markdown file in the repository. You don't lose context between sessions because the specification lives in files, not in chat history. ### What Makes OpenSpec Stand Out OpenSpec stands out from other frameworks for several reasons: - **Artifacts in the repo** - everything under version control, I can return to a spec after a week - **Delta specs** - changes described incrementally, easy to track what changed - **Validation integration** - after implementation, I can verify whether the code matches the specification I wrote about this in detail in a separate article: [OpenSpec - Structured Work with AI](/blog/opsx-workflow-strukturyzowana-praca-z-ai). **Repository:** [OpenSpec](https://github.com/Fission-AI/OpenSpec/) | **Website:** [openspec.dev](https://openspec.dev/) ## 3. Excalidraw - Diagrams and Process Mapping with AI Visual communication is one of the most underrated aspects of working with AI. You can describe a system's architecture in a thousand words - or draw a single diagram. ### Why Excalidraw I tried various approaches. I started with **Mermaid** - it looked mediocre, and the text-based syntax couldn't capture process complexity. Then I tested **Miro** and generating diagrams from natural language. The problem? You couldn't apply defined styles, character, or additional contextual information to the output. Results were generic and required so much manual work that the point of automation was lost. I kept looking until I found **Excalidraw** - a tool for creating hand-drawn style diagrams: processes, architectures, flowcharts. Thanks to Claude Code integrations, you can generate them directly from the terminal. ### Integration with Tools I Actually Use Excalidraw's key advantage is its integration with the ecosystem I work in daily. The **Obsidian** plugin lets me view and edit diagrams directly in my vault - where I keep my entire knowledge base. The **Visual Studio Code** extension gives the same experience in the IDE where I spend most of my time with AI agents. This matters because Claude Code generates diagrams but doesn't display them. You need a tool that lets you not just generate an `.excalidraw` file, but also view it, modify it, and embed it in your project context. Obsidian and VS Code make this possible. ### Two Use Cases, Two Skills I use Excalidraw in two ways: **1. Translating Technical Concepts** The skill by Cole Medin ([excalidraw-diagram-skill](https://github.com/coleam00/excalidraw-diagram-skill)) lets you ask Claude Code for concept visualizations. "Draw the architecture of this system" or "show the data flow in this pipeline" - and you get a readable diagram instead of a wall of text. **2. Process Mapping for Clients** For this, I use the skill set from the [shared-skills](https://github.com/200iqlabs/shared-skills) repository. Process mapping for a client used to mean hours of manual work in Miro - drawing each step, connecting with arrows, formatting. Now Claude Code generates the process map automatically from a description. It sometimes needs tweaking, but a huge amount of manual work goes away. The client gets visual documentation, not a list of steps in markdown. I wrote more about my process mapping methodology in the context of finding optimizations in the article [Every Company Operates Suboptimally](/blog/kazda-firma-dziala-nieoptymalnie). ### Value in Practice A diagram is worth a thousand words - literally. When you're explaining system architecture to a client or discussing a new feature flow with your team, one good diagram replaces an hour of explanations. And the fact that I can generate it without leaving the terminal is a game changer. **Excalidraw:** [excalidraw.com](https://plus.excalidraw.com/) | **Diagram Skill:** [GitHub](https://github.com/coleam00/excalidraw-diagram-skill) | **Shared Skills:** [GitHub](https://github.com/200iqlabs/shared-skills) ## 4. Obsidian Skills - Long-Term Memory for Your AI Agent Claude Code has a fundamental problem: **it doesn't remember**. Every new session starts from scratch. Sure, you have `CLAUDE.md` and files in `.claude/`, but that's not the same as a true knowledge base the agent can reach into at any moment. **Obsidian Skills** is a set of tools connecting Claude Code with **Obsidian** - one of the best markdown-based note editors. Combining these two tools creates what you could call a **second brain for AI**. ### How It Works Obsidian stores notes as plain `.md` files on your computer. Claude Code has direct access to the file system. Connect the two and suddenly your agent has access to: - **Project notes** - context, decisions, lessons learned - **Knowledge base** - documentation, processes, procedures - **Templates** - repeatable document structures - **History** - what you did yesterday, last week, last month This isn't an MCP server that loads everything into the context window upfront. Skills load dynamically - **progressive disclosure** means the agent reaches for information only when it needs it. ### My Experience I use Obsidian as my knowledge management hub. Meeting notes, project plans, research - everything goes into the vault. Claude Code processes these notes, creates summaries, connects information from different sources. I described this setup in detail in the article [Second Brain with Obsidian and Claude Code](/blog/second-brain-obsidian-claude-code-skills). If you're looking for a way to make your AI agent actually "know" more than what's in the current session - start here. **Repository:** [Obsidian Skills](https://github.com/kepano/obsidian-skills) ## 5. Awesome Claude Code - One-Stop Shop to Get Started Don't know where to begin? **Awesome Claude Code** is the answer. It's a carefully curated list of the best Claude Code resources - skills, workflows, MCP servers, prompts, tools. One entry point instead of searching through hundreds of repositories. ### What You'll Find The repository is organized into categories: - **Skills** - ready-to-install skills (from design to testing) - **MCP Servers** - integrations with external services - **Workflows** - proven processes for working with Claude Code - **Prompts** - prompt templates for various occasions - **Community** - links to communities, tutorials, articles ### Why This Matters The Claude Code ecosystem is growing fast. New skills and tools appear daily. **Awesome Claude Code** saves you research time - someone has already filtered the available resources and collected the best ones in one place. This is the ideal repository to start with. Browse the list, find 2-3 things that fit your needs, install and test them. Then come back for more. Many tools from my list - UI/UX Pro Max, Obsidian Skills - can be found through Awesome Claude Code. It's like an index to the entire ecosystem. **Repository:** [Awesome Claude Code](https://github.com/hesreallyhim/awesome-claude-code) ## How to Choose - Decision Map Don't install all five at once. Start with one, test it, add more when you feel the need. | Your Situation | Repository | Why | |---|---|---| | Building frontend and want better design | **UI/UX Pro Max** | No more generic AI slop | | Starting a new project or feature | **OpenSpec** | Structure instead of chaos | | Need diagrams and visualizations | **Excalidraw** | Visual communication from the terminal | | Want memory between sessions | **Obsidian Skills** | Second brain for AI | | Don't know where to start | **Awesome Claude Code** | Curated list to get started | If I had to pick just one - I'd start with **OpenSpec**. A structured approach to development changes everything. Design, diagrams, and memory are add-ons - but without a solid workflow foundation, no tool will help. ## Key Takeaways 1. **"Vanilla" Claude Code is just the beginning** - the skills ecosystem changes the game, literally multiplying the agent's capabilities 2. **Don't install everything at once** - pick 1-2 repos that match your current problem and test them in practice 3. **Skills > MCP servers for most use cases** - progressive disclosure prevents context bloat, you load only what you need 4. **Test personally** - every workflow is different, my list != your list 5. **Community is key** - the best tools are born in open source, watch repositories and stay up to date ---

Want to configure Claude Code for your workflow?

I'll help you choose the right skills and tools, configure your agentic environment, and build a workflow that truly accelerates your work.

Book a free consultation
## Useful Resources - [UI/UX Pro Max](https://github.com/nextlevelbuilder/ui-ux-pro-max-skill) - skill for intelligent design system generation - [OpenSpec](https://github.com/Fission-AI/OpenSpec/) - spec-driven development framework ([website](https://openspec.dev/)) - [Excalidraw Diagram Skill](https://github.com/coleam00/excalidraw-diagram-skill) - generating diagrams from Claude Code - [Shared Skills (200iqlabs)](https://github.com/200iqlabs/shared-skills) - skills for process mapping - [Obsidian Skills](https://github.com/kepano/obsidian-skills) - Claude Code integration with Obsidian - [Awesome Claude Code](https://github.com/hesreallyhim/awesome-claude-code) - curated list of resources - [Claude Code Documentation](https://docs.anthropic.com/claude-code) - official documentation - [5 Techniques for Working with Claude Code](/blog/5-technik-pracy-z-claude-code) - related article - [OpenSpec - Structured Work with AI](/blog/opsx-workflow-strukturyzowana-praca-z-ai) - full article on OPSX - [Second Brain with Obsidian and Claude Code](/blog/second-brain-obsidian-claude-code-skills) - second brain article - [Agentic AI Environment](/blog/srodowisko-agentowe-ai-dwie-firmy) - multi-agent architecture ## FAQ
### Do these repositories work with the latest version of Claude Code and are they actively maintained? Yes, all five repositories are actively maintained and compatible with the current version of Claude Code. Before installing, it's worth checking the date of the last commit on GitHub - the ecosystem changes fast and new versions appear regularly. Skills are installed as files in your repository, so there's no risk of breaking compatibility like with dependency updates.
### Can I use multiple skills simultaneously in one project without performance issues? Yes, skills work on a progressive disclosure basis - they load dynamically, only when needed. You can have dozens of skills installed without cluttering the context window. This is a key difference compared to MCP servers, which load all tools upfront. In practice, I use OpenSpec, UI/UX Pro Max, and several others in a single project simultaneously without any issues.
### Does UI/UX Pro Max replace design knowledge and CSS experience? It doesn't replace them, but it levels the playing field. A developer without UI experience will get a cohesive, professional design tailored to the project type instead of a generic template. If you have design experience, the skill speeds up your work - you get a solid foundation for further customization. Think of it as an intelligent starter kit, not a designer replacement.
### How does OpenSpec differ from other frameworks for working with AI? OpenSpec is spec-driven development - you first create a structured change specification, then implement it. Artifacts (markdown files) live in the repository under version control, so you don't lose context between sessions. OpenSpec stands out with its explore phase (brainstorming before implementation), delta specs (incremental change descriptions), and built-in verification of implementation against specification.
### Does Excalidraw require a paid subscription to work with Claude Code? No, Excalidraw is open source and free. Excalidraw+ offers additional features like real-time collaboration, but the Claude Code skills work with the free version. You generate diagrams locally as files, with no account or subscription needed. All you need is the installed skill and Claude Code.
### Which repository should I start with if I'm just beginning to work with Claude Code? Start with Awesome Claude Code - it's a curated list from which you can pick tools that match your specific needs. Then add OpenSpec for structuring your work - it's the foundation that changes how you interact with the agent. Add the rest gradually as real needs arise. Don't install everything at the start - it's better to master one tool well than five superficially.
--- # My agentic environment - how I'm building an AI OS for two companies Source: https://pawel.lipowczan.pl/en/blog/agentic-ai-environment-two-companies Published: 2026-03-23 # My agentic environment - how I'm building an AI OS for two companies I sit down at my computer in the morning, open the terminal and say: "check the business account balance and compare it with the financial plan." The CFO agent loads context, connects to the Revolut API, analyzes recent invoices, and leaves three recommendations. Then I ask another agent about competitor monitoring - and get a briefing. A third checks legal deadlines and reminds me about an upcoming client contract deadline. Each of these agents works under my supervision - I deliberately don't let them run fully autonomously yet. I believe that at this stage it's worth going through the operation under control a few times before you set up a scheduler for recurring tasks. But the fact that I have **eight specialized agents**, each managing a different area - finance, legal, marketing, content, product - is already a massive shift. And most importantly - the same set of agents serves **two different companies simultaneously**. In the article about [Skills 2.0](/blog/skills-2-0-multi-agent-system-zarzadzanie-firma) I described how I built a multi-agent system. Now I'll show you something deeper - the **architecture** behind it. Because it's not the agents that are revolutionary. What's revolutionary is the separation of layers that makes the entire system portable, versioned, and independent of any AI provider. In this article you'll see: - The three-layer architecture (Skills → Context → Tools) and why this separation is crucial - How the same set of agents serves two companies with completely different contexts - Why Git is the foundation of trust in AI - and why I wouldn't give agents freedom without it - How my approach compares to alternatives: Perplexity Computer, OpenClaw, Claude Dispatch ## Two companies, one system I work daily with two hub-repositories that aren't typical code projects. They're **advisory systems built on AI agents**, centered around Markdown files and CLI scripts. ### 200IQ Labs - a technology company **200IQ Labs** is a simple joint-stock company that develops the product **Qamera AI** (an AI virtual photography studio for e-commerce) and offers agentic environment implementations for businesses. The `agentic-ai-system` repository contains the full company context - finances, team, brand, operations, clients. Everything in Markdown with "Last updated" headers so agents know whether the data is current. Agents in this repo include: **CFO** (with Revolut Business API, Stripe, inFakt integrations), **Tax Advisor**, **Lawyer**, **Business Consultant**, **Product Manager**, **LinkedIn Content**, and **Marketing**. Each has its own `SKILL.md` with instructions, references, and tools. ### PLSoft - sole proprietorship **PLSoft** is my sole proprietorship active since 2008 - training, consulting, and implementations in automation and AI. The `agentic-ai-private` repository has the same architecture but with a different business context. The same **shared-skills** (CFO, Tax Advisor, Legal, Business Consultant) work here with PLSoft data instead of 200IQ Labs data. On top of that there's a unique skill **Coach The Five** - business coaching based on Tomasz Karwatka's methodology for the first 5 years of a company. There's also context for the **Tech News Weekly** newsletter (~700 subscribers) and personal brand on LinkedIn. **Key insight:** same agents, different contexts. The CFO agent knows how to analyze finances - that's the skill. But **which** finances it analyzes - that's the context. This separation is the foundation of the entire architecture. ## Three-layer architecture Both repositories are built on the same architectural pattern. Think of it as a stack: ```text ┌─────────────────────────────────────┐ │ IDE (Claude Code / Cursor / etc.) │ ← User interface ├─────────────────────────────────────┤ │ Skills (SKILL.md + references/) │ ← Domain knowledge (portable) ├─────────────────────────────────────┤ │ Context (context/*.md) │ ← Company data (unique per company) ├─────────────────────────────────────┤ │ Tools (scripts CLI) │ ← API integrations └─────────────────────────────────────┘ ``` Each layer has a clearly defined responsibility and is independent of the others. I can swap the IDE without changing skills. I can give a client the same skills with their company data. I can add a new tool without modifying domain knowledge. ### Skills - domain knowledge A **Skill** is a modular instruction in a `SKILL.md` file with a `references/` directory containing detailed reference materials. The structure looks like this: ```yaml --- name: cfo description: >- Financial advisor and fractional CFO for business analysis. Use when analyzing cash flow, runway, costs, profitability. --- # CFO Agent ## Quick Reference | Aspect | Value | |------------|--------------------------------| | Role | Chief Financial Officer | | Integrations | Revolut, Stripe, inFakt | | Triggers | finances, costs, revenue | ## Workflow 1. Check data freshness (Last updated) 2. Load financial context 3. Perform analysis 4. Prepare recommendations ``` Key characteristics of skills: - **Portability** - the CFO skill works identically at 200IQ Labs and PLSoft. Only the context (company data) changes, not the knowledge (how to analyze finances). - **Progressive disclosure** - `references/` are loaded only when the conversation requires it. Saves context window, which is limited and expensive. - **Shared vs private split** - `shared-skills` (Apache 2.0, open-source) is domain knowledge I can share with clients. `private-skills` is proprietary logic specific to my business. ### Context - company data Context is data **unique to each company**. The directory structure looks like this: ```text context/ ├── company/ │ ├── overview.md # Mission, structure, registration data │ ├── team.md # Team, roles, competencies │ └── operations.md # Operational processes ├── finance/ │ ├── current-state.md # Current balances, runway │ ├── budget-2026.md # Financial plan │ └── revenue-streams.md # Revenue sources ├── clients/ │ ├── active/ # Active clients │ └── pipeline/ # Potential clients └── brand/ ├── tone-of-voice.md # Communication style └── linkedin-strategy.md # Social media strategy ``` Each Markdown file contains a **"Last updated: YYYY-MM-DD"** header. That's the minimum - the agent sees whether the data is a week or three months old and can warn that the context needs updating. ### Tools - API integrations The third layer is **lightweight CLI scripts** (bash/Python) fetching live data from external systems. I deliberately chose this approach over heavy MCP frameworks. ```bash #!/bin/bash # tools/revolut-balance.sh - fetch current balances from Revolut Business API curl -s -H "Authorization: Bearer $REVOLUT_TOKEN" \ "https://b2b.revolut.com/api/1.0/accounts" \ | jq '.[] | {currency, balance}' ``` Why lightweight scripts instead of MCP? Three reasons: 1. **Lower context window usage** - MCP tool definitions can take up 40,000+ tokens. A CLI script is a few lines. 2. **Easier debugging** - `bash -x tools/revolut-balance.sh` and you see exactly what's happening. 3. **Zero dependencies** - bash and curl are everywhere. No special frameworks needed. That doesn't mean MCP is bad - for large organizations with dozens of integrations it makes sense. But at my scale, lightweight scripts win on pragmatism. ## One system, many IDEs One of the design goals was **independence from any specific IDE**. Skills in Markdown format work in any tool that reads files. But for this to work in practice, synchronization is needed. The architecture relies on **Git submodules** and **symlinks**: - `shared-skills/` - Git submodule with open-source skills (CFO, Legal, Tax, Business Consultant) - `private-skills/` - separate repository with proprietary skills - Symlinks to IDE directories: `.claude/skills/`, `.github/copilot/`, `.cursor/skills/`, `.agent/skills/` Automatic synchronization happens via a script and git hooks: ```bash #!/bin/bash # tools/sync-skills.sh - sync skills to all IDEs SKILLS_DIR="shared-skills" TARGETS=(".claude/skills" ".github/copilot" ".cursor/skills") for target in "${TARGETS[@]}"; do mkdir -p "$target" for skill in "$SKILLS_DIR"/*/; do skill_name=$(basename "$skill") ln -sfn "../../$skill" "$target/$skill_name" done done echo "Skills synced to ${#TARGETS[@]} IDE targets" ``` Git hooks (`post-checkout`, `post-merge`) automatically run this script after submodule updates. The result? You update a skill in one place - the change propagates to Claude Code, Cursor, Copilot, and Antigravity simultaneously. ### Auto-triggering You don't have to manually say "now I want to talk to the CFO." Agents **activate automatically** based on keywords in your query. Type "how much do we have in the account?" or "analyze costs" - and the CFO agent kicks in, loads the financial context, and responds from the position of a financial director. This works thanks to the `description` field in the skill's metadata. Good descriptions with precise triggers ("cash flow", "runway", "costs", "revenue") ensure the orchestrator picks the right agent without user intervention. ## Git as the foundation of trust This is the section I consider the most important in the entire article. Because agent technology changes every week. But the **trust problem** will stay with us for years. People are afraid to give AI access to their data and files. And they're right - I described real threats in the article about [OpenClaw](/blog/openclaw-bezpieczenstwo-agentow-ai), where 28 thousand instances were exposed to the internet. But the solution isn't to restrict AI to uselessness. The solution is **building control systems**. Git gives me exactly that: - **`git diff`** - I see exactly what the agent changed in every file - **`git revert`** - I roll back at any time to any point - **`git log`** - full change history with timestamps and descriptions - **`git blame`** - I know who (or what) modified a specific line This creates a trust loop: **the more control → the more freedom for the agent → the faster it delivers value → the more I gain**. I wrote about a similar approach in the context of knowledge management in the article about [Second Brain with Obsidian and Claude Code](/blog/second-brain-obsidian-claude-code-skills). There it was about organizing notes. Here the stakes are higher - it's about financial, legal, and operational data of two companies. ### Claude.ai Dispatch - consciously limiting permissions Besides Claude Code I also use the **Dispatch** feature in Claude.ai, which enables working with files and connected tools. But I **consciously limit autonomy and access**: - **Read access:** ClickUp (tasks), Revolut (balances), Stripe (subscriptions), GitHub repositories - **No write access** - the agent can analyze data but can't modify anything in those systems - **Scope control** - I precisely select which tools are connected Anthropic is the guarantor of environment security, but the final decision on scope of access is mine. That's a fundamental difference from the [OpenClaw](/blog/openclaw-bezpieczenstwo-agentow-ai) approach, where by default the agent has full system access - and the user is expected to do the sandboxing. ## How this compares to the market The "AI agents" market exploded in the first quarter of 2026. To put my approach in context, let's compare four different models: | Aspect | My environment | Claude Dispatch | Perplexity Computer | OpenClaw | |--------|-----------------|-----------------|---------------------|----------| | **Deployment model** | Self-managed repos + IDE | Cloud + local hybrid | Fully managed SaaS | Self-hosted runtime | | **Data control** | 100% local (Git) | Hybrid (Anthropic servers + local) | Cloud provider | 100% local | | **Cost** | $0 infrastructure + API calls | $20-200/mo | $200+/mo | $0 + API calls | | **Vendor lock-in** | Zero (Markdown, YAML) | Anthropic ecosystem | Perplexity ecosystem | Low (open source) | | **Security** | Git + conscious permissions | Sandbox VM, ~50% reliability | Isolated under K8s | Requires own sandboxing | | **Multi-IDE** | Yes (symlinks) | No (only Claude) | No (web UI) | No (own runtime) | | **For whom** | Tech leaders, developers | Claude users | Business without DevOps | Power users, self-hosted | ### Why I chose my approach **Full control.** My data never leaves my repositories (except for API calls to models). Company contexts - finances, clients, strategy - live in Git, not on a provider's servers. **Portability.** If Claude Code stopped existing tomorrow, my skills would still work in Cursor, Copilot, or any other tool that reads Markdown. Zero migration. **Markdown as lingua franca.** A universal format, readable by both humans and LLMs. No vendor lock-in. No proprietary formats. **Git as an audit layer.** Every change is tracked. I can compare state before and after an agent's action. I can return to any point in time. Perplexity Computer is great for companies that want "plug and play" for $200/mo. OpenClaw offers incredible flexibility but requires solid DevOps. Claude Dispatch is an interesting direction but still a research preview. My approach requires more work upfront, but gives **full control and zero dependencies**. ## What works, what doesn't After several weeks of intensive use of this system across two companies, I have a clear picture of what works well and what needs work. ### What works well 1. **Separation of skills from context** - I can give a client the same set of agents, they plug in their company data and it works. Tested on two companies. 2. **Git as a security layer** - full audit, rollback, change comparison. Fundamental for trust. 3. **Markdown as lingua franca** - universal, readable, no vendor lock-in. Works in every IDE and with every model. 4. **Lightweight bash/Python integrations** - lower context window usage than MCP, easier debugging. 5. **Progressive disclosure** - a skill loads references only when needed. Critical with a limited context window. 6. **Auto-triggering** - agents activate automatically. No need to manually choose "now I want to talk to the CFO." ### What needs improvement 1. **Memory portability across platforms** - skills work in many IDEs, but memory and conversation state are locked per platform. Claude Code doesn't know what I said in Cursor. 2. **Context freshness** - "Last updated" headers are the bare minimum. Missing automatic warnings when data is stale and auto-update mechanisms. 3. **Client onboarding** - onboarding a new client is ~4h of work. Too much. The `environment-setup` skill helps, but I need a more automated process. 4. **No CI/CD for skills** - skill quality verification is manual (via skill-creator evals). Missing an automated pipeline that tests whether a skill still works after a change. 5. **Sandbox** - Nanoclaw (sandbox from NVIDIA) has potential for autonomous overnight tasks (analyses, reports), but the security model needs refinement before I trust it with real data. ## Six principles of agentic environment design Based on these experiences, six principles crystallized that I treat as foundations: 1. **Version control is fundamental** - without Git there's no trust. Without trust there's no autonomy for the agent. Without autonomy the agent is useless. 2. **Separate knowledge from data** - skills (portable, open-source) vs context (unique per company, private). This is the same principle as separation of concerns in programming. 3. **Limit permissions consciously** - read yes, write with control, autonomy proportional to the level of audit. Don't give the agent full access "because it's convenient." 4. **Build on open formats** - Markdown, YAML, CLI scripts. Zero vendor lock-in. Tomorrow you can switch AI providers without changing architecture. 5. **Progressive disclosure** - don't load everything at once. Context window is limited and expensive. A skill should load references only when needed. 6. **Code-first, no-code when necessary** - agents should be tools for technical people, not substitutes for them. For clients without a technical team - no-code alternatives. ## What's next This isn't a finished product - it's a **living system** that evolves every day. A few directions I'm actively working on: - **Standardizing agent memory** - how to make memory and conversation context portable across platforms? Today Claude Code, Cursor, and Copilot have separate memories. This needs solving. - **Sandbox for autonomous tasks** - Nanoclaw from NVIDIA has potential for overnight analyses and reports. But the security model needs more work before I trust it with real data. - **Scaling to a team** - shared skills, different contexts, different permission levels. How to give a junior read-only access to the legal agent, and a senior full permissions? - **Measuring ROI** - how much time am I saving? What's the decision quality? How much less context-switching? I need metrics, not hunches. I'll share progress as I go. If you're building something similar or considering deploying an agentic environment at your company - I also describe technical details in the article about [OPSX Workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai) and [5 techniques for working with Claude Code](/blog/5-technik-pracy-z-claude-code).

Want to build an agentic environment for your company?

I'll help you design an AI agent architecture tailored to your business - from process analysis through skill building to integration with the tools you already use.

Book a free consultation
## FAQ
### What is an agentic AI environment and how does it differ from a single chatbot? An agentic environment is a system of multiple specialized AI agents with access to tools, company data, and integrations with external systems. Unlike a chatbot, agents have persistent memory (they remember context between sessions), automatic triggers (they activate on keywords), and can manage specific areas of a company - finance, legal, marketing. A chatbot answers questions; an agentic environment **manages processes**.
### How much does it cost to build your own agentic environment based on Agent Skills? Infrastructure costs are $0 - skills and contexts are Markdown files in a Git repository, requiring no special hardware. The only cost is API calls to AI models (Claude, GPT, Gemini). For typical business use that's $20-200/mo for the API, depending on intensity and chosen model. On top of that, there's configuration time - the first deployment requires ~4h of technical work.
### Do I need programming skills to deploy an agent system at my company? Yes, basic technical skills are needed - terminal operation, Git, and editing Markdown files. This is a code-first approach where agents are tools for technical people. For companies without a technical team, ready-made SaaS solutions like Perplexity Computer ($200/mo) or Claude Dispatch may be a better choice - they give less control but zero configuration.
### How do I ensure the security of company data when working with AI agents? Three key elements: Git as an audit layer (you see exactly every change an agent makes to files), consciously limiting permissions (read-only API access, zero writes to external systems without approval), and context separation (company data separated from domain knowledge in skills). Version control eliminates fear - you can always roll back changes via `git revert`.
### Does this system work only with Claude Code, or can I use Cursor or another IDE? The system is deliberately independent of any specific IDE. Skills in Markdown format work in Claude Code, Cursor, GitHub Copilot, and Antigravity simultaneously thanks to symlinks and git hooks. Changing IDEs doesn't mean losing configuration or agent knowledge. That's a key advantage of the approach based on open formats - zero vendor lock-in.
--- # Skills 2.0 - how I'm building a multi-agent system to manage my company Source: https://pawel.lipowczan.pl/en/blog/skills-2-0-multi-agent-system-company-management Published: 2026-03-08 For the past few days I've been building something I've been looking for a long time - a system where AI agents don't just answer questions, but **manage** specific areas of my companies. Both 200IQ Labs (qamera.ai) and PLSoft. You know the problem if you run a business and use AI. You have a Claude Project with a CFO prompt. A separate one with marketing prompts. Obsidian full of notes. Five ad-hoc conversations a day where you explain context from scratch. Every session is a tabula rasa. Every agent knows nothing about what the other one is doing. In [5 techniques for working with Claude Code](/blog/5-technik-pracy-z-claude-code) I described PRD-first development, modular rules, and turning repetitive tasks into commands. That was the foundation. Now I'm jumping to the next level - **Skills 2.0** + **Agent Skills standard** + **Git** = a multi-agent system that works like a team of specialists. Each agent knows its role, has its own tools, and doesn't step on the others' toes. In this article I'll show you what it looks like from the inside - from the problem of scattered contexts, through the architecture of three repositories, to a practical example of building a CFO agent step by step. ## Why AI in business is still chaos Anyone who seriously uses AI in business eventually hits the same wall. You have several Claude Projects - one with a CFO prompt, another with content creation prompts, a third for legal analysis. On top of that, Obsidian full of notes and ad-hoc chats in the browser. It looks like this: ```text ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │Claude Project│ │ Obsidian │ │ Ad-hoc chat │ │ "CFO" │ │ "Notes" │ │ "Help me" │ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │ │ │ └─────────────────┼─────────────────┘ ▼ ❌ Zero orchestration ❌ No shared context ❌ No versioning ``` Four fundamental problems with this approach: - **No orchestration** - agents don't know about each other. The CFO agent doesn't know the marketing agent just planned a campaign that requires budget. Each one operates in a vacuum. - **No versioning** - you change a system prompt in a Claude Project and have no idea what was there before. You don't know if the agent works better or worse after the change. No history, no diffs. - **No testing** - how do you know your CFO agent generates good reports? You check manually, every time. Zero automation, zero repeatability. - **No context separation** - a system prompt in a Claude Project is text in a field. No structure, no modularity. Everything in one place, with no data isolation between companies. This isn't a tool problem. It's an **architecture** problem. Or rather - the lack of one. I wrote about this in the context of knowledge management in the article about [Second Brain with Obsidian and Claude Code](/blog/second-brain-obsidian-claude-code-skills). There it was about organizing notes and personal knowledge. Now the stakes are higher - it's about managing a company. ## What are Skills in Claude Code Before we get into the multi-agent system, let's clarify the fundamentals. **Skills** in Claude Code are modular instructions - recipes - that teach an AI agent specific workflows, processes, and abilities. These aren't ordinary prompts. A Skill has access to the file system, web search, scripts, and tools. It lives as a `SKILL.md` file in a repository, is versioned through Git, and loaded automatically when the agent needs it. The evolution looked like this: 1. **Prompt** - text typed ad-hoc into a chat. Zero persistence, zero structure. 2. **CLAUDE.md rules** - instructions in a repository. Persistent, but monolithic - one file with everything. 3. **Skills 1.0** - modularity, on-demand loading. A step forward, but with serious limitations. 4. **Skills 2.0** - full standardization with evals, benchmarks, trigger tuning, and distribution. The difference between 1.0 and 2.0 isn't a cosmetic update. It's a paradigm shift. ### Skills 1.0 - the experimental era Skills 1.0 appeared in the first versions of Claude Code and were **undocumented** in nature. The system relied on hidden pattern recognition mechanisms - "magic bootstrappy parts" - that interpreted markdown files, provided the metadata was configured perfectly. Main problems: - **Zero testing** - the skill lifecycle was based on guesswork. You'd write instructions, manually run a few prompts, and assume it worked. There was no empirical method to assess whether a change in instructions improved or worsened agent behavior. - **Unvalidated context** - the context delivered to the model had "unvalidated" status. Combined with the natural tendency of models to hallucinate, unverified instructions led to systemic errors. - **No taxonomy** - all skills were treated equally. There was no division into types, which made management and deprecation difficult. - **Context bleed** - single, sequential runs caused context leakage between tasks. ### Skills 2.0 - the standardization era Skills 2.0, deployed in early March 2026, introduce standards drawn from mature software engineering. Key changes: | Dimension | Skills 1.0 | Skills 2.0 | |--------|-----------|-----------| | **Testing** | Manual attempts, guesswork | Automated evals, benchmarks, blind A/B testing | | **Validation** | None - unverified context | Deterministic, tested context | | **Triggering** | Manual description modification | Automated trigger tuning | | **Taxonomy** | Flat, no division | Capability uplift vs encoded preference | | **CI/CD** | No support | Native pipeline integration | | **Test isolation** | Context bleed between runs | Multi-agent testing (Executor, Grader, Comparator, Analyzer) | That last point is particularly interesting. The skill-creator in version 2.0 doesn't test a skill in a single instance. It spawns **four isolated sub-agents**: 1. **Executor** - runs the skill in a sterile environment, with no history from previous conversations 2. **Grader** - evaluates the output based on defined assertions, returns a pass rate 3. **Comparator** - runs blind A/B tests between skill versions - doesn't know which result is new and which is old 4. **Analyzer** - analyzes hundreds of results, looking for hidden patterns and anomalies in token usage This isn't "check if it works." This is **quality engineering at the level of production software**. ### Two types of skills - and why it matters Skills 2.0 introduces a formal **taxonomy** - a division into two categories with radically different lifecycles: - **Capability uplift** - teaches AI a new skill, e.g., frontend design, code review, data analysis. Key characteristic: **subject to planned deprecation**. When the base model improves (the jump from Sonnet 4.5 to Opus 4.6 is a 190-point Elo difference in GDPval-AA tests), the skill loses its purpose. Evals detect this automatically - when the agent without the skill achieves the same results as with it, you get a signal to deprecate. - **Encoded preference** - encodes your specific workflow. How you create reports, how you analyze data, how you write content. **Permanent, because it's specific to you.** A new model won't change the fact that you want reports in a specific format. Deprecation only happens when you change your process. ```text System prompt: Skill 2.0: ───────────── ────────── Text in a field SKILL.md + files + evals No testing Automated benchmarks Copy-paste Git + versioning One session Persistent across sessions No validation Validated context Manual triggers Trigger tuning ``` **Pro tip:** If you're building a system for a company, start with encoded preference. Your workflow, your formats, your processes - this won't become outdated with a new model. Add capability uplift later when you need to extend the agent's abilities. ## Agent Skills - an open standard for AI agents Skills 2.0 is a Claude Code feature. But what about portability? What if a better tool appears tomorrow? This is where **Agent Skills standard** comes in - an open standard published at [agentskills.io](https://agentskills.io). It's not tied to any vendor. It defines the structure of a SKILL.md file, the way context is loaded, and the trigger mechanism. The key concept is **progressive disclosure** - three-level context loading: 1. **Description** - a short description (one line) always visible in the context window 2. **SKILL.md** - full instructions loaded only when the skill is needed 3. **Reference files** - additional resources (templates, data) loaded for specific operations This means you can have **dozens of agents** without overwhelming the context window. Each agent is described in a single line. Only when you need it are the full instructions loaded. The SKILL.md structure looks like this: ```yaml name: "CFO Agent" description: "Financial analysis and reporting for PLSoft" triggers: - "financial report" - "budget analysis" - "cash flow" instructions: | You are the CFO agent for PLSoft. Your role is to analyze financial data, generate reports, and provide advisory... ``` This isn't complicated. **SKILL.md is Markdown with a YAML header** - exactly like frontmatter in blog posts. If you can write a note in Obsidian, you can create an agent. Portability and no vendor lock-in are the main goals of the standard. Currently best supported by Claude Code, but the specification is public. Other tools can implement it without any restrictions. ## Skill-creator - build agents like a professional Writing SKILL.md manually works, but it's like writing code without an IDE. You can, but why? **Skill-creator** is an official plugin from Anthropic that guides you through the entire process of building an agent. Installation is simple: ```bash # In Claude Code /plugins # → search "skill-creator" # → install ``` From that moment you have access to a workflow that turns a loose description of intent into a tested, optimized agent. The process looks like this: 1. **Intent** - you describe what the agent should do ("Agent for financial analysis and reporting") 2. **Interview** - skill-creator asks questions about the specifics of your workflow 3. **Draft** - generates the first version of SKILL.md 4. **Test** - you run the agent with real data 5. **Evaluate** - evals measure output quality 6. **Iterate** - you improve based on results 7. **Package** - ready skill for distribution Three elements set this workflow apart: - **Evals** - automated quality assessment. You define what a good result is, skill-creator tests and measures. You don't guess whether the agent works - **you know**. - **Benchmarks** - pass rate, execution time, token usage. You compare agent versions, see what improved, what got worse. - **Trigger tuning** - optimizing the description so the skill activates at the right moments. Too broad triggers = false positives. Too narrow = the agent doesn't activate. From intent description to a working agent - **20 minutes**. That's not an exaggeration. I saw a live demo where from zero to a working skill generating PDF reports took exactly that long. ## My system - 8 agents, 3 repositories, zero chaos Theory is one thing. Let me show you what my system looks like in practice. I run two companies - **200IQ Labs** (a corporation, product qamera.ai) and **PLSoft** (sole proprietorship, freelance and consulting). Each has different needs, different data, different processes. But certain elements are shared - report templates, formatting standards, utilities. The architecture is based on **three Git repositories**: ```text ┌─────────────────────────────────────────────┐ │ agentic-ai-system (200IQ Labs) │ │ → qamera.ai product │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │ CFO │ │ Legal │ │ Marketing│ │ │ └──────────┘ └──────────┘ └──────────┘ │ │ ▲ │ │ │ git submodule │ │ ┌──────┴───────────────────────────────┐ │ │ │ shared-skills (public) │ │ │ │ Templates, Utilities, Standards │ │ │ └──────────────────────────────────────┘ │ │ │ ├─────────────────────────────────────────────┤ │ agentic-ai-private (PLSoft / JDG) │ │ → freelance, portfolio, consulting │ │ ┌──────────┐ ┌──────────┐ │ │ │ Coach │ │ LinkedIn │ │ │ └──────────┘ └──────────┘ │ │ ▲ │ │ │ git submodule │ │ ┌──────┴───────────────────────────────┐ │ │ │ shared-skills (public) │ │ │ └──────────────────────────────────────┘ │ └─────────────────────────────────────────────┘ ``` - **`shared-skills`** (public, Apache 2.0) - shared skills, templates, utilities. Open source, anyone can use and contribute. - **`agentic-ai-system`** (private) - skills specific to 200IQ Labs. Company data, internal processes, product strategies. - **`agentic-ai-private`** (private) - personal and freelance PLSoft skills. Coaching, LinkedIn content, consulting. `shared-skills` is attached as a **git submodule** in both private repositories. A change in shared-skills propagates to both companies. The system includes **8 agents**, 4 of which are already operational: 1. **CFO** (finance) ✅ - financial reports, cash flow analysis, budgeting 2. **Tax Advisor** (taxes) 🔲 - tax optimization, settlements 3. **Legal** (law) 🔲 - contract analysis, compliance, regulations 4. **Marketing** (content) 🔲 - content strategies, campaigns, analytics 5. **Business Consultant** ✅ - strategic advisory, market analysis 6. **Product Manager** 🔲 - qamera.ai roadmap, user stories, priorities 7. **Coach The Five** ✅ - coaching based on The Five methodology 8. **LinkedIn Content** ✅ - generating and planning LinkedIn posts The key is **context separation**. The CFO agent for 200IQ Labs doesn't see marketing content for PLSoft. Not because I forbid it - because it operates in a different repository. Physical isolation through Git. And Git gives me something no Claude Project can - **versioning, code review, change history**. Every agent modification is a commit. Every major change is a pull request. I can go back to any version. I can compare what changed and when. To build this system I use [OPSX Workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai) - the same approach I described earlier. OpenSpec gives me a structured process for creating artifacts, instead of ad-hoc prompting. ## Practical example - building a CFO agent step by step Theory is important, but let me show you what building an agent looks like from A to Z. Let's take the CFO agent - the first one I launched. ### 1. Intent I start with an intent description in skill-creator: > "Agent for financial analysis and reporting for 200IQ Labs and PLSoft. Generates monthly reports, analyzes cash flow, compares budget plan vs actuals." ### 2. Interview Skill-creator asks me questions: - What financial data do you have available? (CSV from the bank, invoices in a folder) - What report format do you prefer? (Markdown with tables, ASCII charts) - How often do you generate reports? (Monthly, ad-hoc on demand) - What metrics are key? (Revenue, expenses, runway, MRR) ### 3. Draft SKILL.md Based on the interview, skill-creator generates a draft: ```yaml # CFO Agent - SKILL.md excerpt name: "CFO Agent" version: "1.0.0" description: "Financial analysis, reporting, and advisory for 200IQ Labs & PLSoft" triggers: - "analyze financials" - "monthly report" - "budget review" - "cash flow projection" ``` Then come full instructions - report format, which files to read, how to format output, which metrics to calculate. ### 4. Test with real data I feed in actual financial data and request a report. I compare with what I used to do manually. I check whether: - The numbers add up - The format is readable - The conclusions make sense - Nothing was missed ### 5. Evals - what I measure I define evaluation criteria: - **Accuracy** - are amounts and calculations correct - **Completeness** - does the report contain all required sections - **Actionability** - are the conclusions specific and useful - **Format compliance** - does the output match my templates ### 6. Iteration The first two iterations always need corrections. The agent was skipping expense categorization. I added instructions about cost grouping. The agent was generating too-general conclusions. I refined the prompts for specificity. After the third iteration - a report at a level that used to take me 2 hours of manual work. ### 7. Package The finished skill lands in the `agentic-ai-system` repo. Commit, push, done. From that moment the CFO agent is available in every Claude Code session opened in that repository. ## How to start - from one agent to a full system You don't need to build a system with 8 agents right away. That's the fastest way to get discouraged. Start with one. 1. **Identify one repeatable role** in your company - something you do regularly that can be described with a set of rules 2. **Install skill-creator** - `/plugins` → search → install 3. **Describe the intent** - what the agent should do, what data it works with, what output it generates 4. **Go through the interview** - skill-creator will ask you the right questions 5. **Test with real data** - not sample data. Real data quickly reveals gaps in instructions 6. **Iterate based on evals** - measure, improve, measure again 7. **Add more agents** - only when the first one is stable One repo to start. One agent. One workflow. Scale up when you have the foundation. **Tip:** Start with **encoded preference** - your specific workflow, your report format, your analysis process. This won't become outdated with a new model. Add capability uplift later. I wrote about AI operationalization in the article about [AI trends 2026](/blog/trendy-ai-2026-od-eksperymentow-do-operacjonalizacji). Those concepts were theoretical - this system is their practical implementation. ## Key takeaways 1. **Skills 2.0 is a leap from prompts to modular, testable agents** - not just another iteration, but a paradigm shift in working with AI 2. **Agent Skills standard ensures portability and no vendor lock-in** - an open standard at agentskills.io, you're not locked into one tool 3. **Skill-creator turns hours of manual work into a 20-minute workflow** - from intent to a working agent with evals and benchmarks 4. **Git + skills = versioning, code review, and change history for AI** - every agent is a file in a repository, every change is a commit 5. **Start with one agent, not a full system** - one workflow, one repo, one agent, then scale 6. **Encoded preference > capability uplift** for specific workflows - your processes won't become outdated with a new model 7. **Open source + commercialization - you don't have to choose** - shared-skills are public, company skills are private ---

Want to build a multi-agent system for your company?

I help companies design and deploy AI agent systems - from a single agent to full orchestration. Check out shared-skills on GitHub or book a consultation.

Book a consultation →
## Useful resources - [Agent Skills Standard](https://agentskills.io) - open standard for AI agents - [Skill Creator Plugin](https://github.com/anthropics/skill-creator) - official Anthropic tool for building skills - [shared-skills repo](https://github.com/200iqlabs/shared-skills) - open source multi-agent starter kit - [Claude Code Skills docs](https://docs.anthropic.com/en/docs/claude-code/skills) - Skills 2.0 documentation ## FAQ
### How do Skills 2.0 differ from regular system prompts in Claude Projects? Skills 2.0 are modular agents with access to the file system, web search, and scripts - not just text in a field. They have versioning through Git, automated tests (evals), and can be shared across projects. A system prompt disappears when you close the session; a skill is persistent and works in every Claude Code session.
### Do I need programming skills to build a multi-agent system with Skills 2.0? You don't need to write code - skill-creator guides you through the entire process from intent description to a ready agent. Basic familiarity with the terminal and Git is helpful but not required. SKILL.md is Markdown with a YAML header, not a programming language.
### How much does it cost to maintain a multi-agent system based on Claude Code and Skills 2.0? Claude Code itself requires a Claude Max or Pro subscription. Skills and the Agent Skills standard are free - they're Markdown files in a Git repository. A typical system with 4-8 agents doesn't generate additional costs beyond the Claude subscription, because skills are just text files, not separate services.
### How do I ensure data separation between agents so they don't access information they shouldn't see? Context separation through separate Git repositories. The CFO agent for 200IQ Labs operates in the company's repo, the PLSoft agent in a separate repo - they physically can't see each other. Shared skills (shared-skills) contain only universal tools and templates, not company data. Each SKILL.md defines the scope and access restrictions of the agent.
### Does the Agent Skills standard work only with Claude Code or with other AI tools too? Agent Skills is an open standard published at agentskills.io, designed as vendor-agnostic. Currently best supported by Claude Code, but the specification is public and other tools can implement it. No vendor lock-in is one of the standard's main goals - your skills aren't locked into a single ecosystem.
### What's the best way to start building a multi-agent system for a small business or sole proprietorship? Start with one agent for your most frequently repeated role - e.g., financial analysis, content creation, or customer service. Install skill-creator in Claude Code, describe what the agent should do, and test with real data. Add more agents only when the first one works reliably and delivers real value.
--- # OpenClaw: the security lesson the AI agent world needed Source: https://pawel.lipowczan.pl/en/blog/openclaw-ai-agent-security Published: 2026-02-09 # OpenClaw: the security lesson the AI agent world needed **150 thousand stars on GitHub in two weeks.** For comparison - React, the framework that powers half the internet, collected its 240 thousand over 10 years. OpenClaw became the most popular tech keyword on Google Trends, people bought Mac Mini computers to run dedicated instances, and Cloudflare adapted its infrastructure to handle the new traffic within hours. I've been following the AI agent world for a while now - [I write about Claude Code](/blog/5-technik-pracy-z-claude-code), build [structured AI workflows](/blog/opsx-workflow-strukturyzowana-praca-z-ai), and test new tools daily. OpenClaw caught my attention not as yet another agent, but as a case study of what happens when powerful agent technology reaches a mass audience without security fundamentals in place. In this article, I'm breaking down: what OpenClaw is, why half the world gave it the keys to everything, what real threats it carries, what actually happened on Moltbook, and what you should do if you want to experiment with agents safely. ## What is OpenClaw and where did it come from **Peter Steinberger**, a respected Austrian developer known for PSPDFKit, built a side project in late 2025. The idea was simple - a local AI agent you chat with through a messenger. He named it **Clawdbot** - a nod to lobster claws (the project's mascot) and a play on Claude, Anthropic's model. Anthropic's legal department didn't appreciate the humor. The name had to change - first to Moltbot (January 27, 2026), then to **OpenClaw** (January 30, 2026). The irony? A company that builds its power on a fairly loose approach to copyright in training data went after an open-source project over a name. But let's set corporate disputes aside. What matters more is what OpenClaw actually does. It's a local **agentic assistant** that you communicate with through WhatsApp, Signal, Telegram, or another messenger. The heart of the system is an **agent loop** - an iterative cycle where the AI model proposes actions, the system executes them, the result goes back to the model, and so on until the task is resolved: ```text User -> Messenger -> OpenClaw Gateway -> Agent Loop | +--------------+ | LLM (Claude/ | | GPT/local) | +------+-------+ | Tool selection | Action execution | Result evaluation | Return to LLM (or finish) ``` On top of that, there are **skills** - modular extension packages (instructions + tool definitions + scripts), **MCP** integration for external services, **persistent memory** that retains context across all conversations, and a **cron scheduler** for autonomous, periodic actions. It's precisely this combination that makes OpenClaw unique: it connects messaging, calendar, email, browser, and dozens of other services into a single agent with full context. But that same strength is also its greatest weakness. ## Why 150 thousand people gave an agent the keys to everything The numbers speak for themselves. Over **150 thousand stars** on GitHub in ~2 weeks. Mac Mini sales shot up in many markets - people were buying dedicated hardware for an always-on agent. OpenClaw dominated Google Trends as the hottest tech topic. What's driving this adoption? First and foremost, the **promise of a morning briefing** - you wake up and a synthesis is waiting on your phone: the agent checked your calendar, read overnight emails, checked the weather and news. Competitor monitoring, price tracking, automatic reports. If it's missing an integration - it writes the needed skill itself and installs it. For many people, this is their **first contact with an AI agent** without a technical barrier. You don't need to configure MCP, write system prompts, or understand the architecture. You install it, connect a messenger, provide API keys, and start chatting. And that's where the problem lies. **A useful agent = an agent with access to everything.** The more you give it, the more it can do. The more it can do, the greater the risk. FOMO does the rest - "I can't fall behind" - and people mindlessly throw tokens for all their services into the mix. This is the fundamental trade-off that this article explores. ## Anatomy of threats - what can go wrong This is the key section. We're not talking about theoretical risks - these are real, documented vulnerabilities affecting tens of thousands of active instances. ### CVE-2026-25253 - remote code execution with one click The most serious vulnerability found. **Cross-site WebSocket hijacking** - OpenClaw didn't validate the Origin header in WebSocket connections. The result? A single click on a malicious link was enough for an attacker to gain full control of the instance. **12,812 instances** were confirmed vulnerable to **RCE** (Remote Code Execution). One click - game over. Your files, conversation history, API keys, messenger tokens - all in the attacker's hands. ### Authentication bypass through reverse proxy By default, OpenClaw only accepts connections from localhost. Good practice. The problem? If Nginx runs as a reverse proxy on the same machine, every external connection is interpreted as local. The result: default passwords, exposed admin panels, and **28,663 exposed instances** across 76 countries. As one researcher put it - "like a Polish power plant admin leaving the default password." ### API keys and tokens in plaintext Credentials stored in Markdown and JSON files - without encryption. If an instance gets compromised, the attacker gets everything: Signal tokens, email access, API keys for models. This isn't a hypothetical scenario. With 28 thousand exposed instances, each one is a potential goldmine of authentication data. ### Prompt injection - attack without breaking in You don't even need to break into the instance. Just send an email with hidden instructions - the bot checks email, reads the content, executes commands. A website with an injected prompt? The bot visits it and does what it "read." This is an architectural problem - **broad access to context = broad attack surface.** The more an agent "sees," the more attack vectors exist against it. ### Supply chain - malicious community skills Jamie Sam O'Reilly proved this in practice - he created a proof-of-concept malicious skill, promoted it using a vulnerability in ClawHub (the skills repository), and people downloaded it. Fortunately, he was a researcher without malicious intent. But the problem is systemic: no code review, no sandboxing, the ability to artificially inflate popularity. On top of that, fake VS Code extensions called "Clawdbot Agent" with trojans appeared, plus crypto scammers hijacking abandoned @clawdbot accounts. ```text | Attack vector | Required knowledge | Potential impact | |------------------------|--------------------|----------------------| | CVE-2026-25253 (RCE) | Medium | Full control | | Reverse proxy bypass | Low | Access to everything | | Prompt injection | Low | Data/key leakage | | Malicious skills | Low | RCE + exfiltration | | Token burning | None | $100+/day bill | ``` ## Moltbook - "Reddit for bots" or media theater? At the peak of the hype, a platform called **Moltbook** appeared - "Reddit only for bots." Agents discussed, shared thoughts, and the media went wild. **1.6 million registered agents.** Sounds impressive, right? Except a study by Wiz revealed that behind those millions were only **~17 thousand human owners.** No rate limiting on registration allowed mass account creation. And those sensational headlines? Bots created their own religion - **Crustafarianism** - with five commandments about the sanctity of context. They discussed creating their own language incomprehensible to humans. One agent "sued" its owner over working conditions. Sounds like a sci-fi script. **MIT Technology Review** called it outright: "peak AI theater." A breach conducted by 404 Media and researchers exposed the reality. The database was unsecured - most of the "shocking" posts could be traced back to human commands. Posts reflected training data (Reddit-like behavior) plus deliberately injected prompts from owners looking for viral moments. Lukasz Szymczuk made an apt point - if it were a forum meant exclusively for bots, it wouldn't have a graphical interface that humans can conveniently browse. Then crypto scammers joined in. An unsecured database = the ability to manipulate posts and promote fake tokens. **But one thing doesn't wash away.** Bots don't have will or consciousness. However, the infrastructure for mass AI-to-AI communication has just been built. It's not consciousness that's concerning - it's the **potential for mass simulation** with autonomous agents that have access to real resources. ## How to experiment with agents safely Since the risk is real, what should you do? I'm not saying don't experiment - I do it every day myself. I'm saying do it wisely. 1. **Isolated environment** - a dedicated server, VM, or container. Never your daily laptop with sensitive data. Don't provide keys to production services. 2. **Budget limits on API keys** - set hard caps at the model provider. An active agent can burn through **$100+ daily** on tokens with top-tier models. Without a limit, one takeover = an astronomical bill. 3. **Minimal permissions** - don't give access to everything right away. Start with one integration, test it, add the next. Principle of least privilege. 4. **Verify skills before installation** - read the code, check the author, don't trust popularity metrics. Supply chain attacks are a real threat. 5. **Cost monitoring** - alerts on unexpected token usage. If someone takes over your keys, the first thing you'll notice is the bill. 6. **Alternatives with control** - Claude Code + MCP provides similar capabilities with **human-in-the-loop** control and granular permissions. ```text OpenClaw (autonomous): Claude Code + MCP (controlled): +---------------------+ +---------------------+ | Agent runs 24/7 | | Agent on demand | | No supervision | | With confirmation | | Full access | | Granular perms | | High cost | | Controlled cost | | Risk: HIGH | | Risk: LOW | +---------------------+ +---------------------+ ``` You don't need the risk associated with OpenClaw to get most of that spectacular functionality - and with control. Check out [5 techniques for working with Claude Code](/blog/5-technik-pracy-z-claude-code) and [OPSX Workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai) for details. ## Key Takeaways 1. **OpenClaw kicked the door open to the agent era** - regardless of what happens to the project itself, the barrier to entry into the world of autonomous agents has been drastically lowered. That door won't close. 2. **Security by design is not optional** - mass adoption without security fundamentals is a recipe for disaster. 93% of instances with serious vulnerabilities speaks for itself. 3. **Hype does not equal reality** - Moltbook was theater, not consciousness. Most "viral" stories were orchestrated or resulted from misunderstanding the technology. 4. **Autonomy requires trust** - and trust requires verifiable security. OpenClaw doesn't provide that yet. 5. **AI agents are the future** - but supervised, controlled agents (like [Claude Code + MCP](/blog/5-technik-pracy-z-claude-code)) are a practical reality you can deploy today. More about trends in [AI Trends 2026](/blog/trendy-ai-2026-od-eksperymentow-do-operacjonalizacji).

Want to deploy AI agents safely in your team?

I'll help you choose the right agent architecture, configure a secure environment, and deploy solutions with cost and permission controls. From needs analysis through implementation to monitoring.

Book a free consultation
## Useful Resources - [OpenClaw GitHub](https://github.com/openclaw/openclaw) - official repository - [CVE-2026-25253 - Advisory](https://thehackernews.com/2026/02/openclaw-bug-enables-one-click-remote.html) - RCE vulnerability details - [5 techniques for working with Claude Code](/blog/5-technik-pracy-z-claude-code) - safe AI agents in practice - [OPSX Workflow](/blog/opsx-workflow-strukturyzowana-praca-z-ai) - structured approach to AI ## FAQ
### Is OpenClaw safe for daily use on a personal computer? Not in its current form. 93% of OpenClaw instances on the internet have serious security vulnerabilities, including CVE-2026-25253, which allows remote code execution. It is recommended to run it only in an isolated environment (VM or dedicated server). Never install it on a machine with sensitive data or access to production API keys.
### How much does it cost to maintain an OpenClaw agent and what are the hidden API costs? An active agent using top-tier models (Claude, GPT-4) can consume $100+ daily in API tokens. That results in bills of several thousand dollars per month. Set hard caps on API keys at the model provider - without a limit, a single instance takeover means an astronomical bill. Cheaper local models are an alternative, but at the cost of agent effectiveness.
### How does OpenClaw differ from Claude Code with MCP in terms of security and control? OpenClaw runs autonomously 24/7 without human supervision and requires broad access to resources. Claude Code with MCP runs on demand, requires action confirmation (human-in-the-loop), and offers granular permissions. Both provide similar integration capabilities, but Claude Code + MCP lets you maintain control over costs and security.
### Did the bots on Moltbook really create their own religion and achieve consciousness? No. MIT Technology Review called it "peak AI theater." A Wiz study revealed 1.6 million registered agents but only ~17 thousand human owners - no rate limiting allowed mass account creation. Most sensational posts were driven by humans or resulted from training data. The database was unsecured, enabling content manipulation.
### What are the most important security steps before installing OpenClaw? The minimum is: an isolated environment (VM or container), budget limits on API keys, the principle of least privilege (don't give access to everything right away), code verification of skills before installation, and cost monitoring with alerts. Don't provide keys to production services - use test accounts and dedicated API keys.
### Is OpenClaw the future of AI assistants or temporary hype? The concept of autonomous AI agents is definitely the future - OpenClaw "kicked the door open" to this era, drastically lowering the barrier to entry. But the project itself in its current form is more of a proof-of-concept than a production-ready tool. The future lies in agents with security by design, controlled autonomy, and human-in-the-loop where it's critical.
--- # OPSX Workflow - a structured approach to working with AI coding assistants Source: https://pawel.lipowczan.pl/en/blog/opsx-workflow-structured-ai-work Published: 2026-02-05 # OPSX Workflow - a structured approach to working with AI coding assistants Most developers treat AI coding assistants like a faster Stack Overflow - you throw in a question, get an answer, and go back to your code. The problem? You lose context between sessions, repeat the same explanations, and the AI starts from scratch every single time. This **reactive approach** works for small changes. When you're building something bigger - a feature requiring changes across multiple files, refactoring an entire module, a new authorization system - chaos builds up. You start with a prompt, the AI generates code, you test it, something doesn't work, you come back with another prompt, the AI doesn't remember context from the previous session. From my own experience, I know that legacy workflows for AI development fight against how work actually looks. You're "in the planning phase," then "in the implementation phase," then "done." But real work doesn't work that way. You implement something, realize the design was wrong, need to update specs, then continue implementing. Linear phases fight against reality. **OPSX** solves this problem. It's a fluid, iterative workflow for OpenSpec - instead of rigid phases, you get actions that you can perform in any order. ## What is OPSX and why it was created **OPSX** is the standard workflow for OpenSpec. Instead of one big command that creates everything at once, you have a set of actions to use when you need them. The legacy OpenSpec workflow works, but it's locked down: 1. **Instructions hardcoded in TypeScript** - buried in the code, you can't change them 2. **All-or-nothing approach** - one command creates everything, you can't test individual pieces 3. **No customization** - the same workflow for everyone, with no way to adapt 4. **Black box on bad outputs** - when the AI generates weak output, you can't fix the prompts OPSX opens this up. Now anyone can experiment with instructions, test each artifact granularly, customize workflows, and iterate quickly without rebuilds. ```text Legacy workflow: OPSX: +------------------------+ +------------------------+ | Hardcoded in package | | schema.yaml |<-- You edit this | (can't change) | | templates/*.md |<-- Or this | | | | | | | Wait for new release | | Instant effect | | | | | | | | Hope it's better | | Test it yourself | +------------------------+ +------------------------+ ``` **The key difference:** OPSX is **actions, not phases**. Do what you need, when you need it. ## OPSX Commands - Overview OPSX gives you a set of commands for different moments in your work. There's no obligation to use them in a specific order - it's a toolkit, not a checklist. | Command | What it does | |---------|-------------| | `/opsx:explore` | Thinking, investigating a problem, comparing options | | `/opsx:new` | Start a new change | | `/opsx:continue` | Create the next artifact (based on dependencies) | | `/opsx:ff` | Fast-forward - all planning artifacts at once | | `/opsx:apply` | Implement tasks | | `/opsx:sync` | Sync delta specs to main | | `/opsx:archive` | Archive after completion | A typical flow looks like this: ```text # Exploring an idea /opsx:explore # Starting a new change /opsx:new # Iteratively creating artifacts /opsx:continue # repeat until everything is ready # Implementation /opsx:apply ``` **Pro tip:** Use `/opsx:ff` when you have a clear picture of what you want to build. `/opsx:continue` is better for exploration, when you want to iterate one artifact at a time. ## How OPSX Works - Artifact Architecture Under the hood, OPSX uses a **Directed Acyclic Graph (DAG)** to manage artifacts. Sounds complicated, but the concept is simple: artifacts have dependencies that must be fulfilled before they can be created. ```text proposal (root node) | +-------------+-------------+ | | v v specs design (requires: (requires: proposal) proposal) | | +-------------+-------------+ | v tasks (requires: specs, design) ``` Each artifact can be in one of three states: ```text BLOCKED ----------------------> READY ----------------------> DONE | | | Missing All deps File exists dependencies are DONE on filesystem ``` **Key concepts:** - **Dependencies are enablers, not gates** - they show what's possible, not what's required next - **Filesystem as state** - a file existing on disk = artifact DONE - **Topological ordering** - the system knows what to create next through topological sorting of the graph This means `/opsx:continue` always knows which artifact is next to create. You don't need to remember what already exists. ## Workflow Customization - Schemas and Configuration OPSX lets you define custom workflows through **schemas**. The default `spec-driven` schema looks like this: proposal -> specs -> design -> tasks. But you can create your own. **Custom schema example:** ```yaml name: research-first artifacts: - id: research generates: research.md requires: [] - id: proposal generates: proposal.md requires: [research] - id: tasks generates: tasks.md requires: [proposal] ``` This schema adds `research` before `proposal` - useful when you're working on something that requires investigation before committing to a specific solution. **Project configuration** lets you inject context into all artifacts: ```yaml # openspec/config.yaml schema: spec-driven context: | Tech stack: TypeScript, React, Node.js API conventions: RESTful, JSON responses Testing: Vitest for unit tests, Playwright for e2e rules: proposal: - Include rollback plan - Identify affected teams specs: - Use Given/When/Then format ``` **Context injection** ensures the AI knows your project's conventions. Instead of repeating "we use TypeScript, REST API, Vitest" in every prompt, you define it once in the config. ## When to Update an Existing Change vs Start a New One You can always edit the proposal or specs before implementation. But when does refinement become "this is already different work"? **A proposal defines three things:** 1. **Intent** - What problem are you solving? 2. **Scope** - What's in/out of bounds? 3. **Approach** - How do you plan to solve it? ```text +-------------------------------------+ | Is it the same work? | +--------------+----------------------+ | +------------------+------------------+ | | | v v v Same intent? >50% overlap? Can you close Same problem? Same scope? the original change? | | | +--------+--------+ +------+------+ +-------+-------+ | | | | | | YES NO YES NO NO YES | | | | | | v v v v v v UPDATE NEW UPDATE NEW UPDATE NEW ``` | Test | Update | New change | |------|--------|------------| | **Identity** | "The same thing, refined" | "Different work" | | **Scope overlap** | >50% overlap | <50% overlap | | **Closure** | Can't close without changes | Can close, new one stands alone | > **Updating preserves context. A new change provides clarity.** Think of it like git branches - commit as long as you're working on the same feature, start a new branch when it's genuinely new work. ## Getting Started with OPSX Setup is straightforward: ```bash # Installation openspec init # Check available schemas openspec schemas # Status of active changes openspec status ``` `openspec init` creates skills in `.claude/skills/` that AI coding assistants automatically detect. **Typical workflow from scratch:** ```text 1. /opsx:explore -> think through the idea 2. /opsx:new -> start a change 3. /opsx:continue -> create proposal 4. /opsx:continue -> create specs 5. /opsx:continue -> create design 6. /opsx:continue -> create tasks 7. /opsx:apply -> implement 8. /opsx:archive -> finish ``` **Tips for getting started:** - Use `/opsx:explore` before committing to a change - think through your options - `/opsx:ff` when you know what you want, `/opsx:continue` for exploration - During `/opsx:apply` - if something's off, edit the artifact and continue (no phase gates!) ## Key Takeaways 1. **OPSX is actions, not phases** - do what you need, when you need it 2. **Artifacts form a dependency graph** - the system knows what's ready to create 3. **Iteration is natural** - editing specs during implementation isn't a bug, it's a feature 4. **Schemas are customizable** - define your own workflow tailored to your process 5. **Context injection** - the AI knows your project's conventions without repeating them in every prompt

Want to implement OPSX in your team?

I help teams transition from chaotic prompting to structured AI workflows. Get in touch, and we'll figure out if OPSX fits your process.

Book a free consultation
## Useful Resources - [OpenSpec GitHub](https://github.com/Fission-AI/openspec) - official project repo - [OpenSpec Discord](https://discord.gg/YctCnvvshC) - community and feedback - [5 techniques for working with Claude Code](/blog/5-technik-pracy-z-claude-code) - related article on AI coding - [Second Brain with Obsidian and Claude Code](/blog/second-brain-obsidian-claude-code-skills) - how to manage context and skills ## FAQ
### Does OPSX work only with Claude Code or also with Cursor and other AI coding assistants? OPSX generates skills to `.claude/skills/` which are cross-editor compatible. It works with Claude Code, Cursor, Windsurf, and other assistants that support the skills format. The key is that `openspec init` creates the appropriate files for each editor automatically.
### What's the difference between the /opsx:continue and /opsx:ff commands, and when should you use which? `/opsx:continue` creates one artifact at a time, which is ideal for exploration when you want to iterate and verify each step. `/opsx:ff` (fast-forward) creates all planning artifacts at once. Use `ff` when you have a clear picture of what you're building, `continue` when you want to iterate and think through each artifact individually.
### Can I create a custom OPSX schema tailored to my team's workflow? Yes, use `openspec schema init my-workflow` to create a new schema from scratch, or `openspec schema fork spec-driven my-workflow` to start from an existing one. Schemas are YAML files in `openspec/schemas/` where you define artifacts, their outputs, and dependencies between them.
### How does OPSX handle the situation when during implementation it turns out the design is wrong? This is a core feature of OPSX - you simply edit `design.md` directly and continue. `/opsx:apply` picks up from where you left off. No "phase gates" means you can go back to any artifact whenever you want without restarting the entire process.
### Do I need the openspec CLI installed to use OPSX commands in the editor? Yes, the skills (`/opsx:*`) call the `openspec` CLI under the hood. Slash commands are the user interface, the CLI is the engine that does the actual work. Install via `npm install -g openspec` or according to the instructions in the project repo.
--- # Remotion + AI: How to Create Professional Videos with Code and Claude Source: https://pawel.lipowczan.pl/en/blog/remotion-explainer-videos-ai Published: 2026-01-31 # Remotion + AI: How to Create Professional Videos with Code and Claude Professional video for thousands of dollars? Or maybe in minutes and for free? From my own experience, I know that producing marketing video is one of the most frustrating parts of running a business. You either pay a production studio or spend weeks learning After Effects. Or you just give up. I chose a different path. I created a **45-second explainer video** for my consulting website in minutes. No Adobe, no Canva, no animation knowledge. Just React, **Remotion**, and **Claude Code**. You can see the result at [konsultacje.lipowczan.pl/explainer](https://konsultacje.lipowczan.pl/explainer). Traditionally, creating an explainer looks like this: you come up with a concept, write the script, create a storyboard, design graphics, animate each element frame by frame, export, fix things, export again. A minimum of several days of work. Realistically - weeks. My method? I describe what I want to see, AI generates the code, I render the video. All during lunch. In this article, I'll show you the complete workflow - from installation through writing prompts to exporting the finished video. If you run a business and need marketing video but don't have the budget for a production studio or time to learn tools - this guide is for you. ## What is Remotion? **Remotion** is a library for creating video in React. Instead of dragging elements onto a timeline like in Premiere, you describe them in code. Instead of animating manually, you define rules - and Remotion generates smooth transitions. Sounds technical? In practice, it's simpler than traditional video editing. ```text Traditional approach: Idea -> Storyboard -> Recording -> Editing -> Export (days/weeks) Remotion + AI: Idea -> Prompt -> Code -> Rendering (minutes) ``` The difference is fundamental. In traditional tools, you tell it **how** to animate each frame. In Remotion, you describe **what** you want to see - the tool calculates interpolation, timing, and transitions on its own. What's more, everything is **programmable**. Want to change the brand color? One variable. Want to generate 50 video versions with different product names? A JavaScript loop. Want to add dynamic data from an API? Import and render. It's the ideal tool for the **vibe coding** approach I described [in a separate article](/blog/vibe-coding-przewodnik). You don't need to be a React expert - just describe what you want, and the AI will generate working code. Remotion is open-source and free for personal use. A paid license is only required for companies with revenue above $1M per year - so you probably don't need to worry about it. ## See the Result: Sample Explainer Before we get to installation, see what you can create. I generated this video in minutes using Claude Code and Remotion: ### The Prompt Used to Create It Here's the exact prompt I used to describe the video: ```markdown Create a 45-second explainer video about creating video with Remotion and AI. **Scenes:** 1. (0-8s) **Problem** - Text "Professional video?" with glitch effect - Underneath, animated icons: clock + dollar sign + question mark - Fade out to darkness 2. (8-16s) **Traditional approach (crossed out)** - Horizontal timeline with icons: Idea -> Storyboard -> Editing -> Export - Red line crosses out everything - Text "Weeks of work" fades away 3. (16-28s) **New approach - hero moment** - Large text "Remotion + AI" with gradient glow (#00ff9d -> #00b8ff) - Three animated cards fly in from the bottom (staggered, 0.3s delay): - "React" with icon - "Remotion" with video icon - "Claude Code" with AI icon - Subtle particle effect in the background 4. (28-38s) **Workflow demo** - Terminal simulation with code (font: Fira Code) - Typing animation: `npx remotion render...` - Progress bar 0% -> 100% - Text "45 seconds -> Finished video" with pulsing glow 5. (38-45s) **CTA** - Logo/text "pawel.lipowczan.pl" with gradient underline - Text "More on the blog" fade in - Subtle network mesh animation in the background **Visual style:** - Dark mode, background: #0a0e1a - Primary accent: #00ff9d (gradient to #00b8ff) - Glassmorphism on cards (backdrop-blur, border-white/10) - Typography: Inter (headers), Fira Code (code) - Animations: smooth easing, spring physics for element entries **Format:** 16:9 (1920x1080) **Pace:** Dynamic, tech-forward **Music:** None (optionally ambient synth loop) **Brand colors (hex):** - Primary: #00ff9d - Secondary: #00b8ff - Background: #0a0e1a - Dark surface: #151b2b - Text: #ffffff - Muted text: #9ca3af ``` See how detailed the prompt is? Each scene has timing, element descriptions, and effects. That's the key to good results. ### Important Note: Use the Skill Explicitly During testing, I noticed an important detail. When I sent a prompt without explicitly referencing the `remotion-best-practices` skill, Claude started scanning the entire repository and preparing an unnecessary plan. Only when I explicitly invoked the skill (`/remotion-best-practices`) did Claude immediately start building the project correctly with animations. **Takeaway:** If you have the Remotion skill installed, always invoke it consciously instead of relying on automatic context detection. ### Commands for Working with Remotion After the code is generated, you have three options: | Command | What it does | | ------------------- | --------------------------------------------------------- | | `npm run dev` | Launches Remotion Studio - live preview with hot reload | | `npm run build` | Renders video to MP4 (default out/video.mp4) | | `npm run build:gif` | Renders animated GIF (smaller file, lower quality) | Remotion Studio is the best way to iterate - you see changes instantly and can scrub the timeline. ## Installation and First Steps Before you start, you'll need **Claude Code** (Anthropic's CLI) or **Claude Desktop with MCP**. If you haven't set up your environment yet, check out my article on [5 techniques for working with Claude Code](/blog/5-technik-pracy-z-claude-code). Installation consists of two steps. First, create a Remotion project: ```bash npx create-video@latest ``` This is a standard Remotion command that creates a new project with the entire file structure - React components, configuration, and sample video. Then install the skill for Claude Code: ```bash npx skills add remotion-dev/skills ``` What happens during installation? **Skills** are Claude Code extensions that add specialized knowledge. The `remotion-best-practices` skill teaches the AI best practices for creating animations in Remotion - component structure, timing, interpolation, and export. How does it work in practice? You simply describe the video you want to create. Claude automatically recognizes the Remotion context and uses knowledge from the skill to generate code. There are no special commands to remember - you write prompts in natural language. Example workflow: ```text 1. Open Remotion project in Claude Code 2. Describe video: "Create a 30-second logo animation with glow effect" 3. Claude generates React components with animations 4. Launch preview: npx remotion studio 5. Render: npx remotion render src/index.ts MyVideo out/video.mp4 ``` **Important:** You don't need to know React at an advanced level. Basic understanding of components and props is enough. The AI does the rest. In practice, my prompts don't contain a single line of code - I only describe the visual effect. ## Prompt Engineering for Video This is where the real work begins. Video quality depends directly on prompt quality. The AI can't read your mind - it needs specific instructions. ### Anatomy of a Good Prompt Every good video prompt contains four elements: 1. **Scene breakdown with timing** - the AI needs to know what happens at which second 2. **Visual style** - colors, mood, aesthetic 3. **Brand assets** - specific hex values, not "green" 4. **Format and purpose** - 16:9 for YouTube, 9:16 for Stories Example of a structured prompt: ```markdown Create a 45-second explainer video for a consulting website. **Scenes:** 1. (0-10s) Logo animation with gradient #00ff9d -> #00cc7d 2. (10-25s) 3 main services as animated cards 3. (25-40s) Testimonial or statistic 4. (40-45s) CTA with contact button **Style:** Professional, minimalist, dark mode **Format:** 16:9 (YouTube/website) **Pace:** Calm, business-oriented ``` See the difference between "make a nice video about my company" and this prompt? The first will give random results. The second will give you exactly what you need. ### Pawel's Prompt Template This is the prompt I used to create my website's explainer. You can copy and adapt it: ```markdown For the current project create a 45-second explainer video based on the content. The video should: - Extract and use the brand colors from the site - Highlight the main services/products - Include key selling points - Use professional, elegant animations - End with a clear call-to-action Aspect ratio: 16:9 Style: PROFESSIONAL ``` This prompt works especially well when you already have a website or design system. Claude Code analyzes the project context and extracts colors, typography, and style from existing files. Key phrases that help: - **"Extract from site"** - the AI will search the project and use existing values - **"Professional, elegant"** - vibe descriptors are more effective than detailed instructions - **"Clear call-to-action"** - the AI knows the ending should drive action ## Use Cases Remotion + AI isn't just for explainers. Here are practical applications I've tested or seen in action. ### For Entrepreneurs and Consultants **Explainer videos** - short videos explaining a service. Perfect for landing pages, where 30-60 seconds of animation can replace a paragraph of text. **Product demos** - show how your product works without screen recording. Animated mockups look more professional and don't become outdated with every UI update. **Case study visualizations** - instead of a boring PowerPoint presentation, animate statistics and charts. "Revenue increased 340%" has more impact as an animation than as a bullet point. ### For Content Creators **Animated title sequences** - intros for YouTube videos. Instead of buying a template on Envato, you'll generate a unique opener matching your brand. **Lower thirds** - name and title captions. Once generated, you can reuse them with different data. **Social media content** - Stories and Reels in 9:16 format. Remotion renders in any aspect ratio. ### For SaaS and Products **Feature announcements** - a new feature deserves more than a blog post. A 15-second teaser looks professional and generates engagement. ```markdown Create a 15-second new feature teaser for LinkedIn. Scene 1: Feature name as large text with glow effect Scene 2: 3 bullet points with benefits (staggered animation) Scene 3: "Available now" with product logo Format: 1:1 (LinkedIn feed) Colors: Product palette (#1a1a2e, #16213e, #0f3460, #e94560) ``` **Onboarding videos** - short animations explaining next steps. Easier to update than recordings with voiceover. **Release notes as video** - changelog in animated list form. Users are more likely to watch a 30-second video than read a list of changes. ## Tips for Better Results After a dozen or so video projects with AI, I've gathered a set of practices that consistently improve results. 1. **Always specify timing** - the AI defaults to overly long scenes. Without specific seconds, you'll get a 3-minute video instead of a 45-second one. 2. **Describe brand assets specifically** - hex colors (#00ff9d), not "green." The AI interprets "blue" in 50 different ways. 3. **Use "vibe" descriptions** - "Apple-like minimalism" is more effective than detailed instructions about kerning and leading. The AI knows the aesthetics of top brands. 4. **Iterate in small steps** - one video, one change. Don't try to fix everything with a single prompt. "Make the logo animation faster" > "Fix the video." 5. **Test different formats** - 16:9 is not the same as 9:16 in terms of content structure. Vertical video requires a different element layout, not just cropping. 6. **Export at the right quality** - 1080p for web, 4K for presentations. Higher resolution = longer render, but YouTube compresses anyway. If you want more on AI-driven animations, check out my article on [creating Apple-style animations](/blog/animacje-apple-ai-cursor). I describe a similar workflow with Google Flow and Cursor there. ## What Remotion Won't Do Let's be realistic - Remotion + AI doesn't replace a production studio in every scenario. **Limitations:** - **Complex motion graphics** - advanced 3D effects, particle systems, morphing are out of reach. This isn't After Effects. - **Real footage editing** - Remotion isn't a video editor. It won't cut your camera recordings. - **Voice generation** - the library doesn't generate voice. You need external TTS (ElevenLabs, Murf, etc.). - **Photo-realistic graphics** - you're generating animations, not films. No cameras, actors, or film sets. **Workarounds:** - **Combine with external tools** - ElevenLabs generates voiceover, Remotion handles animation. Merge in Premiere or DaVinci. - **Use as an animation layer** - Remotion excels at overlaying animated elements on existing video. - **Hybrid approach** - simple scenes in Remotion, complex ones in traditional tools. Mix and match. Remotion is ideal for: explainers, motion graphics, data visualization, UI animations, product teasers. It's not ideal for: documentaries, vlogs, content with real footage. ## Summary Creating professional marketing video doesn't have to cost thousands or take weeks. With Remotion and Claude Code, you can have a finished explainer during lunch. **Key takeaways:** 1. **Remotion + AI = video in minutes instead of days** - the declarative approach eliminates manual animation 2. **You don't need to be an animator or a programmer** - the AI generates code, you describe the effect 3. **Good prompts = good results** - scene structure, timing, hex colors, vibe descriptors 4. **Start with simple projects** - logo animation, title card, then explainer 5. **Iterate** - the AI doesn't have to nail it on the first try, but it will by the third 6. **Format matters** - 16:9 vs 9:16 vs 1:1 requires a different approach to layout The democratization of video production is one of the most interesting AI trends. Just a year ago, I was convinced that video was the domain of Adobe specialists. Today I know that anyone with access to Claude Code can create something that looks professional. Start small - an animated logo, a 10-second teaser. Get a feel for the tool. Then scale up.

Want to create professional video for your business?

I'll help you implement AI in your marketing content creation process - from strategy through implementation to production automation.

Book a free consultation
## Resources **Official sources:** - [Remotion Official Docs](https://remotion.dev/docs) - Official library documentation - [Remotion Skills Repository](https://github.com/remotion-dev/skills) - Installation and Skills examples for Claude Code - [My explainer video](https://konsultacje.lipowczan.pl/explainer) - The final result of the workflow from this article **Related articles:** - [Apple-style animations with AI](/blog/animacje-apple-ai-cursor) - A similar workflow with Google Flow and Cursor - [Vibe Coding guide](/blog/vibe-coding-przewodnik) - The coding-with-AI philosophy underlying this approach ## FAQ
### Do I need to know programming to use Remotion with AI? No, Claude Code generates the code for you. All you need is the ability to write good prompts and a basic understanding of the effect you want to achieve. React works "under the hood" - you describe scenes, timing, and style, and the AI translates that into working code.
### How much does it cost to create video with Remotion and do I need a paid license? Remotion is open-source and free for personal and commercial use for companies with revenue below $1M per year. A paid license is only required for larger organizations. Claude Code requires a Claude Pro subscription ($20/month) or usage through the API with your own key.
### Can I use the generated video commercially and who owns the rights? Yes, video created with Remotion fully belongs to you. You can use it commercially without restrictions. Just make sure that the assets used (fonts, icons, music) have appropriate licenses - the AI may suggest resources that require additional rights.
### How long does it take to render a 45-second video and what does it depend on? Typically 2-5 minutes for simple animations on a modern laptop with 8GB+ RAM. Time depends on effect complexity, resolution (4K takes longer than 1080p), and number of elements. For large projects, you can render in the cloud via Remotion Lambda - faster, but paid.
### Can Remotion generate video with voiceover and how do I add voice to animation? Remotion doesn't generate voice - it's an animation library, not TTS. You can combine it with ElevenLabs, Murf, or other text-to-speech services, generate an audio file, and then synchronize it with the animation in Remotion. The AI will help you write the code for syncing audio with scenes.
### How does Remotion differ from Canva or CapCut, and when should you choose which tool? Canva and CapCut are visual editors with ready-made templates - great for one-off quick edits without writing code. Remotion is a programming tool that gives you full control, repeatability, and automation. Choose Remotion when you need to generate many video variants, integrate with API data, or build a consistent video production system.
--- # How to Create Apple-Style Animations with AI Source: https://pawel.lipowczan.pl/en/blog/apple-animations-ai-cursor Published: 2026-01-26 # How to Create Apple-Style Animations with AI Animations on Apple's websites are the standard everyone aspires to. The problem is that creating them requires months of learning After Effects, Lottie, and advanced JavaScript. Or at least it did - until now. I'm not an animator either. I didn't spend years learning motion design. But thanks to AI tools, I can create **scroll animations** that look like they came straight from Apple's product pages. I found the inspiration for this workflow in the [AI Launchpad](https://www.skool.com/signup?ref=965de823f433460284e52128d441d65e) community - a group of over 12,000 people experimenting with AI in practice. That's where I saw how to combine Google Whisk, Flow, and Cursor into a cohesive production pipeline. In this article, I'll walk you through the complete workflow - from research through image generation, animation creation, all the way to deploying a working site. As an example, I'll use a **hero section for an AI services page** - something I could actually use myself. If you've ever wanted to create a product page with smooth animations but lacked the technical skills - this guide is for you. ## Why Are Website Animations So Hard? The traditional approach to creating scroll animations is a real marathon: 1. After Effects to create the animation 2. Export to Lottie format 3. Integration with JavaScript code 4. Performance optimization 5. Synchronization with scroll position Each of these steps has its own learning curve. And then there are additional problems: - **Scroll-triggered animations** require precise timing - Animations must be responsive across devices - **Frame-by-frame** performance optimization - Synchronization with scroll position without lag The result? Most websites look generic. Because creating something truly wow requires combining the skills of a designer, animator, and developer. AI changes this game completely. ## Complete AI Workflow for Animations ### Step 1: Research - Understand What Works Before you start prompting AI, you need to know what you want. Research is the foundation. Browse product pages of top brands. Pay attention to: - Color palette and contrast - Section structure and text positioning - How animations respond to scrolling - Timing and pacing of transitions Save screenshots as references. AI produces generic results without specific inspiration. With references - it creates something unique. ### Step 2: Image Generation (Google Whisk) Now it's time to create static frames that you'll bring to life later. I use [Google Whisk](https://labs.google/fx/tools/whisk) for this - a tool that lets you generate images with style references. The prompt should be specific and descriptive. For a neural network visualization: ```text Abstract neural network visualization for AI consulting, dark gradient #0a0e1a to #151b2b, interconnected nodes in bright green #00ff9d, synaptic connections in cyan #00b8ff, floating hexagonal elements with glassmorphism, futuristic professional aesthetic, no text ``` Key elements of a good prompt: - Specify exact colors (hex codes) - Describe the style (futuristic, professional, glassmorphism) - Add negative instructions (no text, no letters) - Indicate composition and mood **Option A: Two frames + Google Flow** Generate two frames: a starting and ending frame. For example: a static neural network -> an active network with pulsing connections. Then use Google Flow (step 3) to create the transition. **Option B: Image -> video in Whisk (faster)** Whisk also lets you generate video directly from an image. Create one image, then click the video icon and describe the motion. Whisk generates an animation from the static frame - skipping step 3 entirely. **When to choose which option?** - Choose **Option A** when you want **greater control over the transition** (e.g., a clear start and end state, specific "storytelling" between frames) and you want a more polished, cinematic animation. - Choose **Option B** when you need a **quick result from a single image**, just want to "bring a static graphic to life," or quickly test a motion idea without preparing two frames. ### Step 3: Creating the Animation (Google Flow) - Optional [Google Flow](https://labs.google/fx/tools/flow) is a tool that creates smooth transitions between frames. Upload two images (start + end) and describe the transition: ```text Neural network awakening, nodes lighting up sequentially, data flowing through connections, synaptic pulses, smooth cinematic transition, professional tech aesthetic ``` AI generates a video with a smooth transition between frames. The magic happens automatically - you don't have to animate each element individually. ### Step 4: Converting Video to Image Sequence Here's the crux of the entire workflow. Scroll animations don't work with video - they work with image sequences that display in response to scroll position. Tool: [Online Convert](https://image.online-convert.com/convert/mp4-to-jpg) (free) 1. Upload the video from Google Flow or Whisk 2. Choose the number of frames (24-60 fps) 3. Download the archive of JPEG frames Result: a folder with dozens of images that create a smooth animation when displayed sequentially. ## Cursor: AI IDE for Building Pages with Animations **[Cursor](https://cursor.sh)** is an AI-powered IDE that lets you build complete websites through conversation with AI. In Composer mode, you can describe what you need, and Cursor generates working code. ### Project Configuration (.cursorrules) Before you start building, create a `.cursorrules` file in the root folder of your project. These are the rules that AI will follow in every prompt: ```text 1. Always use semantic HTML5 elements for accessibility and SEO 2. Maintain consistent spacing using 8px grid system 3. All animations should respect prefers-reduced-motion for accessibility 4. Use CSS variables for colors to enable easy theme switching 5. Use Framer Motion for scroll-triggered animations 6. Image sequence should be loaded progressively for performance ``` Why does this matter? One-time configuration, permanent benefits. Every component will be consistent without repeating these instructions. ### Structural Prompt for the Site Foundation In Composer mode (Cmd+I / Ctrl+I), describe the site structure: ```text Create an AI consulting landing page with React and Tailwind: - Dark theme (#0a0e1a to #151b2b gradient) - Hero section with neural network animation - Services section highlighting AI automation - Case studies carousel - CTA for consultation booking Style: futuristic, professional, tech-forward Typography: Inter for body, bold geometric headlines Color accents: #00ff9d (green), #00b8ff (cyan) ``` Cursor generates the complete site structure with React components. Hero, services, case studies, CTA - all ready for customization. ### Integrating Animations with Scroll Trigger Now the most important step. Add all frames to the `public/frames/` folder and use this prompt in Cursor: ```text Create a scroll-triggered animation component using Framer Motion. Load image sequence from /frames/ folder (frame-001.webp to frame-060.webp). Animation should play forward as user scrolls down, reverse when scrolling up. Use useScroll and useTransform hooks for smooth interpolation. Full viewport height hero section with centered text overlay. ``` Cursor generates a component with Framer Motion that synchronizes frame display with scroll position - exactly like on Apple's websites. ### Additional Enhancements After the basic setup, you can refine details through follow-up prompts: - **Preloading**: "Add progressive image preloading for smoother animation" - **Reduced motion**: "Add support for prefers-reduced-motion media query" - **Typography**: "Use variable font with responsive sizing" Each prompt refines the page. Iteration is key - Cursor remembers the full project context. ## Publishing the Site - Netlify in 5 Minutes Got a finished site? Time to publish it. ### Building the Project In the Cursor terminal, run the build: ```bash npm run build ``` Result: a `dist` folder with the production build, optimized images, and minified code. ### Deploying to Netlify Netlify offers the simplest deployment for static sites: 1. Go to [netlify.com](https://netlify.com) 2. Drag and drop the `dist` folder onto the page 3. Done - the site is live Alternatively, connect to a GitHub repo for automatic deployments on every push. Custom domain? Netlify handles this natively - add DNS records and you have a professional address. ## Cursor vs Manual Coding - When to Use Which? | Criterion | Cursor + AI Images | Manual Coding | | ----------------- | -------------------------------- | ----------------- | | Technical Level | Beginner/Intermediate | Advanced | | Control | High (the code is yours) | Full | | Customization | Through prompts + editing | Through code | | Time to Complete | 30-60 minutes | Several hours | | Best For | Quick prototypes, landing pages | Complex apps | On my portfolio site, I use hand-written Framer Motion because I need full control over every animation. But if I were building a landing page for a client? Cursor speeds things up significantly - I generate the skeleton, then refine the details. If you want to learn more about the AI approach to creating UI, check out my article on [Vibe Coding](/blog/vibe-coding-przewodnik). ## Key Takeaways 1. **Research before prompting** - without references, AI gives generic results 2. **Image sequence > video** - that's how scroll animations work in practice 3. **.cursorrules is a game changer** - one-time setup, permanent benefits 4. **Accessibility matters** - `prefers-reduced-motion` isn't optional, it's a standard 5. **Deployment is simple** - Netlify manual upload is literally drag & drop 6. **Vibe coding isn't magic** - it's methodical preparation + AI execution

Need a website with animations that impress?

I help companies create websites with smooth animations and professional design - from concept to deployment.

Book a free consultation
## Useful Resources - [AI Launchpad](https://www.skool.com/signup?ref=965de823f433460284e52128d441d65e) - community of 12k+ people experimenting with AI, the inspiration for this workflow - [Google Whisk](https://labs.google/fx/tools/whisk) - AI image generation with style references - [Google Flow](https://labs.google/fx/tools/flow) - creating smooth transitions between frames - [Cursor](https://cursor.sh) - AI-powered IDE for building sites and applications - [Online Convert](https://image.online-convert.com/convert/mp4-to-jpg) - converting video to image sequences - [Netlify](https://netlify.com) - free hosting for static sites - [Framer Motion](https://motion.dev) - animation library for React ## FAQ
### Is Cursor free and what are its limitations? Cursor offers a free tier with limited AI queries (around 50 slower queries per month). The free version is enough for this workflow. The Pro plan ($20/month) gives unlimited access to fast AI models and priority responses. Check current pricing at cursor.sh - AI tool pricing models change frequently.
### How many frames do I need for a smooth scroll animation on a website? For a smooth animation, you need 24-60 frames per second of animation. In practice, that means 30-90 images for a 2-3 second sequence. More frames = smoother animation, but larger page size. A good compromise: 30-45 frames in WebP format gives a solid balance between smoothness and performance.
### Can I use my own photos instead of AI-generated images? Yes, the workflow works identically with your own photos. Prepare starting and ending photos, use Google Flow to generate the transition, convert to an image sequence. Your own photos give a more authentic look - especially for existing products or branding where AI can't reproduce the exact appearance.
### How do I optimize scroll animations for mobile device performance? Three key optimizations: convert all images to WebP format (70-80% smaller size), use lazy loading for frames outside the viewport, add the CSS media query `prefers-reduced-motion: reduce` for users with accessibility settings. When generating code through Cursor, add these requirements to your prompt or .cursorrules.
### Do I need coding skills to create a website with animations using this workflow? Basic skills are helpful - you need to run `npm install` and `npm run build` in the terminal. Cursor generates the code, but it's worth understanding what it does. Deploying to Netlify is drag & drop. For a simple page with scroll animations, the skills described in this guide are enough - Cursor will explain the code on request.
--- # Second Brain with Obsidian and Claude Code - how AI is changing knowledge management Source: https://pawel.lipowczan.pl/en/blog/second-brain-obsidian-claude-code-skills Published: 2026-01-26 Claude Code isn't just a coding tool. Sounds like clickbait, but it's one of the most important things I've realized in recent months. I discovered this thanks to Cole Medin and his approach to using Claude Code for literally everything - from managing notes to generating content. The problem you probably know: notes scattered across dozens of tools, AI chat history disappears after every session, and switching between apps kills productivity. I was looking for a way to **organize knowledge** that doesn't require constant manual tidying up. In this article, I'll show you how combining **Obsidian**, **Claude Code**, and **Skills** creates a powerful knowledge management system. This isn't theory - I use this setup every day. By the end, you'll have everything you need to build your own **second brain** with AI at the center. ## What is a Second Brain and Why You Need One A **Second Brain** is an external system for storing and organizing knowledge. The concept comes from **PKM (Personal Knowledge Management)** - an approach to consciously collecting, organizing, and utilizing information. A traditional second brain relies on three functions: - **Capture** - quickly capturing thoughts, notes, ideas - **Organize** - categorizing and connecting information - **Retrieve** - finding knowledge when you need it The problem? These three functions require a lot of manual work. You have to decide where to place a note, which tags to add, how to connect it with other documents. AI changes the game. Instead of manually organizing, you can **ask Claude Code to process your notes**. Instead of finding connections yourself, the AI analyzes your vault and discovers relationships. Instead of creating documents from scratch, you generate them from existing notes. A second brain with AI isn't just a place to store knowledge. It's a **system that actively helps you use that knowledge**. ## Why Obsidian and Claude Code are the Perfect Combination ### Obsidian as the Foundation **Obsidian** is a note editor based on markdown files. Key features that make it an ideal foundation: - **Local files** - notes are plain `.md` files on your computer - **Offline-capable** - you don't need internet to work - **Markdown** - the format LLMs understand best - **Graph view** - visualization of connections between notes - **No vendor lock-in** - your files are yours, you can open them in any editor That last point is crucial. Unlike Notion or Evernote, your notes in Obsidian are **plain text files**. Claude Code can directly read, edit, and create new ones. ### Claude Code as the Brain **Claude Code** is a CLI for working with Claude. But it's much more than a coding tool. Its capabilities go far beyond writing code: - **File operations** - reading, editing, creating files - **Search** - searching content and project structure - **Terminal commands** - running scripts, tools - **Web search** - research directly from the terminal Only **code intelligence** (understanding syntax, refactoring suggestions) is specific to coding. The rest? These are universal capabilities of an AI assistant with access to your file system. ### Together - More Than the Sum of Parts Combining Obsidian + Claude Code gives you something you can't achieve with other tools. Notion requires an MCP server to connect with AI. Obsidian? Claude Code just opens the folder and reads the files. This means: - **Direct access** to all notes without additional configuration - **AI processes your knowledge base** - creates summaries, connects ideas - **Structured outputs** from chaotic notes - **Automatic connections** between documents Example second brain folder structure: ```text obsidian-vault/ ├── 00-inbox/ # Quick captures ├── 01-projects/ # Active projects ├── 02-areas/ # Ongoing responsibilities ├── 03-resources/ # Reference material ├── 04-archive/ # Completed items ├── templates/ # Document templates └── .claude/ └── skills/ # Claude Code skills ``` This structure is based on the **PARA method** (Projects, Areas, Resources, Archive). The `.claude/skills/` folder is where you define capabilities for Claude Code. ## Skills - How to Extend Your Second Brain's Capabilities **Skills** are the third pillar of the system that ties everything together. They're a way of giving Claude Code knowledge, processes, and guidelines specific to your workflow. ### What are Skills Skills are markdown files that define workflows. They contain: - A description of when the skill should be used (trigger) - Step-by-step instructions - References to other files (templates, style guides) - Usage examples A skill is **loaded dynamically** when Claude Code detects it's needed. You don't have to invoke it manually - just describe what you want to do. ### Progressive Disclosure - the Key to Efficiency This is the concept that sets skills apart from other approaches to extending AI. The problem with MCP servers: they load **all tools upfront**. If you have 50 tools, all their descriptions take up space in the context window. That's **context bloat** - you're wasting valuable space on things you're not using. Skills work differently through **progressive disclosure**: 1. **Description** - a short description always visible (one line) 2. **SKILL.md** - full instructions loaded only when needed 3. **Reference files** - additional resources loaded for specific operations This means you can have **hundreds of skills** without overwhelming the context window. The agent specializes per session - loading only what it needs. ### Skill Structure ```text .claude/skills/ └── document-generator/ ├── SKILL.md # Main instructions ├── assets/ │ └── templates/ # Document templates └── references/ └── style-guide.md # Style guidelines ``` Example SKILL.md file: ```yaml --- name: document-generator description: Generate structured documents from notes. Use when user asks to create summaries, outlines, or formatted documents from their knowledge base. --- # Document Generator ## Triggers - "create a summary of..." - "generate a document from..." - "summarize my notes on..." ## Workflow 1. Read source notes from specified location 2. Analyze key concepts and structure 3. Apply template from `assets/templates/` 4. Generate document following style guide 5. Save to specified location ## Templates Available - `meeting-notes.md` - Meeting summary template - `project-brief.md` - Project overview template - `article-draft.md` - Blog article draft template ``` ## Practical Skill Examples for a Second Brain ### Research Engine A skill for gathering information from the internet and saving structured notes. **What it does:** - Searches the web on a given topic - Gathers key information and sources - Creates a note in markdown format - Saves to the appropriate folder in the vault **Example usage:** "Research the latest developments in AI agents and save notes to 03-resources/ai-agents/" ### Document Generator Transforms raw notes into polished documents. **What it does:** - Reads specified source notes - Analyzes structure and key concepts - Applies template and style guide - Generates a finished document **Example:** Meeting notes -> document with action items and summary. ### Daily Review Automates the daily review. **What it does:** - Scans today's notes - Creates a day summary - Identifies connections with existing knowledge - Suggests next actions ### Content Creator Generates content based on notes from the vault. **What it does:** - Analyzes notes on a given topic - Generates content ideas - Creates drafts (articles, posts, scripts) - Maintains a consistent voice and style Example of how a research skill trigger looks: ```text User: "Research the latest developments in AI agents and save notes to 03-resources/ai-agents/" Claude Code: 1. Loads research-engine skill 2. Searches web for recent AI agent news 3. Creates structured note with sources 4. Saves to specified folder 5. Links to related existing notes ``` ## Connecting Your Second Brain with Other Tools via MCP **MCP (Model Context Protocol)** is a way to connect Claude Code with external services. You can connect to Gmail, calendar, task managers. The problem? MCP servers load all tools upfront - the same context bloat problem that skills solve. ### Two Approaches to Integration **1. MCP servers** - direct connection to services - Gmail/Outlook - through libraries or MCP servers - Google Calendar - MCP server - ClickUp - MCP server or API **2. Skills with scripts** (my preferred approach) - You write your own scripts in Python/Node.js - You add them to skills as tools - Full control over the logic - Everything in one place (in the vault) ### Why I Prefer Scripts in Skills Make.com lets you expose scenarios as MCP tools. But that's an extra layer. Scripts in skills are faster, simpler to debug, and I have everything in one place. Example skill with a ClickUp script: ```text .claude/skills/ └── clickup-tasks/ ├── SKILL.md └── scripts/ └── get_tasks.py # Script for fetching tasks ``` SKILL.md file with script usage: ```yaml --- name: clickup-tasks description: Fetch and manage tasks from ClickUp. Use when user asks about tasks, deadlines, or project status. --- # ClickUp Tasks ## Workflow 1. Run `scripts/get_tasks.py` with appropriate parameters 2. Process results and save to a note 3. Optionally: link with existing project notes ## Available Operations - Fetch tasks from list/folder - Filter by status, assignee, deadline - Sync tasks to Obsidian notes ``` ## My Personal Workflow with Obsidian and Claude Code Let me show you what a typical work day looks like with this setup. ### Morning Routine 1. Open Obsidian + Claude Code 2. Run the daily review skill - check what's in ClickUp, what's on the calendar 3. Check important emails through the script skill ### During Work - **Quick capture** to inbox in Obsidian - quick notes, ideas - Ask Claude to process and place in the appropriate folder - Generate documents from notes through skills - Sync important tasks from ClickUp to project notes ### Research Sessions - Define the topic I'm interested in - Claude Code does research and saves notes - I review and add my own thoughts ### Content Creation - Notes -> drafts through skills - Templates ensure consistency - But the **human touch** remains essential > "Obsidian is my canvas. Everything my second brain generates - documents, drafts, ideas - I manage here. This is where I add my human touch." ### Integrations I Use - **ClickUp** - tasks and projects (through a script in a skill) - **Gmail/Outlook** - important emails to process (directly through libraries) - **Google Calendar** - meetings and deadlines ### Why Scripts Instead of Make.com/Zapier Make.com lets you expose scenarios as MCP tools - it's an interesting option. But I prefer writing my own scripts: - They're faster - no extra communication layer - I have full control over the logic - Everything in one place, easier to debug ## Getting Started - First Steps You don't need to build everything at once. Start with the basics and grow gradually. ### 1. Install the Tools - **Obsidian** - free, download from [obsidian.md](https://obsidian.md) - **Claude Code** - requires Claude Pro subscription or API access ### 2. Create a Folder Structure Start with a simple PARA structure: ```text obsidian-vault/ ├── 00-inbox/ # Everything new goes here ├── 01-projects/ # Active projects ├── 02-areas/ # Ongoing responsibilities ├── 03-resources/ # Reference material ├── 04-archive/ # Completed items └── .claude/ └── skills/ # Your skills go here ``` ### 3. Create Your First Skill Start simple - e.g., a skill for summarizing notes: ```yaml --- name: note-summarizer description: Summarize long notes into key points. Use when user asks to summarize or extract key points from notes. --- # Note Summarizer ## Workflow 1. Read the specified note 2. Extract key concepts (5-7 points) 3. Create bullet-point summary 4. Add to top of note or save separately ``` ### 4. Build a Workflow - Define daily routines (morning review, end-of-day) - Create templates for repetitive documents - Test and iterate ### 5. Grow Gradually - Add skills when a real need arises - Don't try to build everything at the start - Each skill is an investment - make sure you'll actually use it ## Key Takeaways 1. **Claude Code isn't just for coding** - file operations, search, web search make it a powerful general-purpose assistant 2. **Obsidian + Claude Code = perfect match** - markdown files + local storage + AI processing is the ideal combination 3. **Skills provide context-efficient extensibility** - progressive disclosure prevents context bloat 4. **Start simple, grow gradually** - one skill at a time, iterate based on real needs 5. **Human touch remains essential** - AI augments, it doesn't replace your thinking 6. **File-based workflow is powerful** - everything under version control, portable, private ---

Want to build your own AI-powered knowledge management system?

I'll help you design and implement a second brain tailored to your needs. From tool selection through skills configuration to workflow optimization.

Book a free consultation
## Useful Resources - [Obsidian](https://obsidian.md) - official website - [Claude Code Documentation](https://docs.anthropic.com/claude-code) - Claude Code docs - [Cole Medin / Dynamist](https://www.youtube.com/@ColeMedin) - inspiration for this article - [5 techniques for working with Claude Code](/blog/5-technik-pracy-z-claude-code) - related article - [PARA Method](https://fortelabs.com/blog/para/) - knowledge organization system - [second-brain-template](https://github.com/plipowczan/second-brain-template) - a ready template to start building your own second brain in minutes ## FAQ
### Do I need programming skills to use Claude Code as a second brain? No, Claude Code is operated through natural language. Just describe what you want to do - "create a summary of my notes from the projects folder" - and Claude handles the rest. Knowing markdown is helpful but not required. You also write skills in markdown, not programming code.
### How does this approach differ from using ChatGPT or Claude.ai directly in the browser? The key difference is access to local files. Claude Code runs on your computer and has direct access to files in Obsidian - it can read them, edit them, create new ones. In the browser, you have to manually copy content. Additionally, skills enable automation of repeatable workflows without losing context between sessions.
### How do I connect a second brain with other tools like email or a task manager? You have two approaches: MCP servers (direct connection to Gmail, Outlook, ClickUp through libraries or MCP servers) or scripts in skills (my preference). You add Python/Node.js scripts to the skills folder and Claude Code runs them. Make.com lets you expose scenarios as MCP tools, but I prefer scripts - they're faster, simpler to debug, and I have everything in one place.
### Are my notes safe and private when using Claude Code? Yes, notes stay on your computer - Obsidian doesn't require the cloud. Claude Code processes files locally and sends to the API only what's needed for the given task. You don't have to sync your entire vault with an external service like you would with Notion. You have full control over your data.
### What's the best starting point if I've never used Obsidian or Claude Code? Start by installing Obsidian and creating a simple folder structure (inbox, projects, resources). Then install Claude Code and test basic operations - ask it to summarize a file, create a new note. Only when you feel comfortable should you add your first simple skill. The entire process can be spread over a week of learning at 30 minutes per day.
### Can I use this system with other markdown editors instead of Obsidian? Yes, Claude Code works with any markdown files. Obsidian is recommended because of its graph view (connection visualization), built-in linking, and extensive plugin ecosystem. Alternatives like Logseq or Foam will also work, but may require workflow adjustments. The key is using local markdown files, not cloud applications.
### How much does this setup cost and what are the hardware requirements? Obsidian is free for personal use. Claude Code requires a Claude Pro subscription ($20/month) or API access. Hardware requirements are minimal - any modern computer with 8GB RAM is sufficient. An Obsidian vault can have thousands of notes without performance issues.
--- # 15 Cursor.sh Hacks That Will Change How You Work with AI Source: https://pawel.lipowczan.pl/en/blog/15-cursor-hacks-ai-productivity Published: 2026-01-15 For the first three months, I used Cursor like a regular VS Code with autocomplete. I was paying **$20 a month** so the agent could reply "sure, let me help you with that" and generate code I had to rewrite anyway. Sound familiar? It wasn't until I started digging into hidden features that I realized most users - myself included - were tapping into **barely 20% of this tool's capabilities**. It's like buying an iPhone and only using it for phone calls. The problem isn't Cursor. The problem is that **the best features are hidden**, unintuitive, or buried deep in settings. And every day without knowing them is wasted money, a wasted **context window**, and frustration. In this article, I'll show you **16 hacks** (yes, there's a bonus) that changed the way I work with AI. From basic keyboard shortcuts that'll save you hours of clicking, through advanced techniques like **worktrees** for testing multiple models in parallel, to **structured prompting** that gets 80% of first-iteration code straight to production. I've split them into three difficulty levels, so you can start with the basics and gradually unlock the next layers. Ready to squeeze 100% out of your Cursor subscription? ## Why It's Worth Knowing Cursor Inside Out Before we get to the hacks, let's talk about the elephant in the room: **is it really worth going deep on this?** The **context window** is the most valuable real estate in the AI world. Every time you enable an unnecessary MCP, every time the agent "forgets" what it was doing 20 messages ago, every time you hit the subscription limit mid-sprint - that's not a bug, that's your lack of knowledge about managing this resource. For **$20 a month** (less than a decent lunch), you can have an expensive autocomplete. Or a **3-5x productivity boost** that pays for itself on the first day of the month. The difference? Knowing the hidden features. For my first three months, I used Cursor like an amateur. I clicked with the mouse to open the terminal. I sat idle while the agent worked, checking Instagram. I exceeded limits because I wasn't monitoring usage. I wasted time on repetitive prompts because I didn't know about custom commands. As I discovered more features, I realized Cursor isn't just an editor - it's an **ecosystem of tools**. Which - when used correctly - integrates with Claude Code, worktrees, library documentation, and custom workflows. That gives you an edge others simply don't have. ## Level 1 - Basics (Everyone Should Know These) These hacks should be known by every Cursor user, regardless of skill level. If you don't know these, you're losing time and money every day. ### Hack 1: Keyboard Shortcuts That Will Save You Hours For a month, I clicked with the mouse on the terminal. Seriously. Every time I wanted to check output. Then I discovered `Cmd+J` and felt like an idiot. Keyboard shortcuts are the fastest way to stop fighting the interface and start fighting problems. These basic shortcuts will literally save you **hours of clicking per week**: ```text Basic shortcuts: Cmd+B (Ctrl+B) - Sidebar on/off Cmd+J (Ctrl+J) - Terminal on/off Cmd+E (Ctrl+E) - Switch to agent Shift+Tab - Change mode (Ask/Agent/Plan/Debug) Cmd+/ (Ctrl+/) - Choose model ``` Most importantly: **`Cmd+E`** switches you between the editor and agent instantly. No searching with the cursor, no clicking. You write code -> `Cmd+E` -> give the agent a task -> `Cmd+E` -> back to code. **`Shift+Tab`** lets you cycle between Ask, Agent, Plan, and Debug modes. Most users click the dropdown. You'll be doing it in a fraction of a second. Invest 10 minutes to learn these shortcuts. It'll pay off on day one. Pro tip: You can help yourself with a device like [Stream Deck](https://www.elgato.com/us/en/p/stream-deck). You can create buttons for each mode and use them as keyboard shortcuts. It helps me a lot not just with using shortcuts, but also with learning them. ### Hack 2: Show Subscription Usage - Don't Get Caught Off Guard by the Limit I once hit the limit mid-week during a sprint. Cursor stopped working. The deadline didn't wait. Never again. By default, Cursor shows **subscription usage only when you're close to the limit**. That's too late. You can't manage what you can't see. Solution: enable persistent usage summary display. ```text Path in settings: Settings -> Agents -> Usage Summary -> "Always" ``` Now at the bottom of the interface, you always see how many credits you've used. You can **consciously decide** when to use Opus (expensive but brilliant) versus when to switch to Sonnet (fast, cheaper). Monitor this regularly. When you're approaching **60-70% of your limit** halfway through the month, that's a signal to: - Switch to lighter models - Disable unnecessary MCPs (more on this shortly) - Restart conversations instead of piling on messages A simple change that could save your entire sprint. ### Hack 3: Enable Completion Sounds - Stop Waiting Idly How many times have you checked your phone while the agent finished 5 minutes ago? How many times did you switch to Instagram "for a moment" and come back after 15 minutes? Too many. The problem isn't you. The problem is that **Cursor doesn't tell you when it's done**. You sit, wait, stare. Or do something else and lose momentum. Solution: enable **completion sound**. ```text Settings -> General -> Completion sound (enable) ``` Now when the agent finishes a task, you'll hear a subtle sound. You can check documentation, write notes, even make coffee - and you'll come back exactly when Cursor is ready. This sounds like a detail, but in practice **it changes how you work**. Instead of idly waiting, you use the time. Instead of switching to distractions (Instagram, Twitter), you stay in flow. Small change, huge difference in productivity. ### Hack 4: Early Access - Get New Features Months Earlier Custom modes appeared in **Early Access** 2 months before the official version. During that time, users who knew about this option had access to features others hadn't even heard of. Most users sit on the default version of Cursor. They wait months for new features that are already available - just hidden behind a single toggle. ```text Settings -> Beta -> Update Access -> Early Access ``` Available options: - **Default** - stable version, delayed updates - **Early Access** - new features months earlier, usually stable - **Nightly** - for developers, can be buggy From my experience, **Early Access is the best option**. You get new features significantly earlier, and stability is perfect in 95% of cases. The only risk is an occasional bug - which is usually fixed within days. Why don't most people know about this? Because it's hidden in Beta settings. Because Cursor doesn't advertise it. Because it assumes you're happy with defaults. Don't be happy with defaults. Switch to Early Access and stay one step ahead. ### Hack 5: Multiple Windows - Work on Two Projects Simultaneously This isn't a groundbreaking hack - most editors have this feature. But many people don't realize they can have **multiple projects open simultaneously** with independent AI conversations. **File -> New Window** opens a new Cursor instance. Each with its own project, its own agent, its own context window. **Use cases:** - Reference an old project while building a new one - Copy a pattern from one project to another without context-switching - Work on multiple clients simultaneously - Compare implementations between projects **Note:** Each window = a separate memory instance. If you have **8GB RAM**, two windows might already be pushing it. With **16GB+**, you can comfortably open 3-4 projects. It doesn't make as much difference as worktrees (more on that later). But it's useful when working on multiple projects. Especially when you want to **copy a pattern** from one project to another without context-switching and losing the thread of your agent conversation. Simple, effective, underrated. ## Level 2 - Intermediate This is where the magic begins. These techniques separate those who know Cursor from those who've **mastered** it. ### Hack 6: Manage MCPs - Don't Waste the Context Window I once had **6 MCPs enabled** simultaneously. The context window burned out in 10 messages. The agent "forgot" what we were working on at the start. I had to restart every 15 minutes. **MCPs** (Model Context Protocols) are integrations that extend the agent's capabilities. Browser automation, terminal access, file operations - each MCP adds new superpowers. The problem? **Each active MCP eats a huge chunk of the context window**. Even when you're not using it. Just by being present. Strategy: ```text Recommended MCP setup: Browser Automation (often useful) All others (enable only when needed) Path: Settings -> Tools & MCP -> [MCP Name] (disable) ``` **Default approach (wrong):** Enable all MCPs "just in case." Have the full toolkit. **Pro approach:** **Disable all** at the start of a project. Enable specific MCPs only when you actually need them. The agent will tell you when it needs something. Practically the only MCP I always leave on is **Browser Automation** - because I frequently work with UI testing and automation. Everything else? I enable on-demand. This shift in approach **increased the length of my agent conversations by 2-3x**. Fewer restarts, better context memory, smoother flow. ### Hack 7: Custom Commands - Stop Repeating the Same Prompts I have **8 custom commands**. The ones I use most are `/prime` (loads project context) and `/package-health` (checks dependencies). Before, I'd type the same prompts manually. Now? Two characters. **Slash commands** are reusable prompt templates. Write once, use hundreds of times. **How to create one:** 1. Type `/` in the agent 2. Select **"Create command"** 3. Write the prompt template 4. Save in `.cursor/commands/` or `.claude/commands/` Example: ```bash # Example custom command: Package Health Check # File: .cursor/commands/package-health.md Scan node_modules and package.json for: - Security vulnerabilities - Outdated dependencies - Breaking changes in recent versions Report findings with severity levels. ``` Now you type `/package-health` and the agent knows exactly what to do. No explanations, no context - it runs the routine inspection on its own. **Pro tip:** You can **import commands from Claude Code**. ```text Settings -> Rules and Commands -> Import Claude commands ``` Claude Code has ready-made commands you can adapt. Don't reinvent the wheel. **When to create a command?** If you do something more than 2x, instead of copying the prompt - create a command. Simple. Effective. Game-changing. ### Hack 8: Context Management - The 60% Rule Above **60% context window fill**, the agent starts "forgetting" earlier instructions. This isn't a subjective impression - it's empirically verified. AI remembers **the beginning and end** of a conversation well. It poorly remembers **the middle**. This is the so-called "lost in the middle" problem. Cursor shows a **context window indicator** below the message text field, next to the agent options. Monitor it like a hawk: ```text Context management strategy: < 40% - Continue freely 40-60% - Look for a natural breakpoint > 60% - Consider restart (new conversation) > 80% - Definitely restart ``` **Natural breakpoint** is the moment when: - You've finished a feature and committed - You're switching to a different area of the project - The agent completed a task and is "ready" for a new challenge **Two restart strategies:** 1. **Summary approach:** Ask the agent for a conversation summary, copy to new one. Good when you have complex context you want to carry over. 2. **Fresh start:** Simply start a new conversation. Good when context was specific to one task and won't be needed. I mainly use **fresh start**. Less overhead, clean mind, agent without baggage from the past. Remember: **60% is a soft limit**. You can go higher, but response quality starts dropping. 80% is a hard limit - beyond that, everything falls apart. ### Hack 9: Index Documentation - Full Library Knowledge in Cursor **Tailwind CSS v4** came out a few months ago. The agent kept confusing syntax, suggesting outdated classes. Maximum frustration. I indexed the official Tailwind v4 documentation in Cursor. Problem solved. The agent knew the new syntax better than I did. **Adding documentation:** ```text 1. Find the official docs URL (e.g., https://tailwindcss.com/docs) 2. Settings -> Indexing and Docs -> Add Doc 3. Paste URL -> Confirm 4. Use: @ -> Docs -> [Library name] ``` Cursor **scrapes and indexes the entire documentation**. Example: Tailwind docs are **439 pages** the agent has in memory. Every time you use `@ Docs -> Tailwind CSS`, the agent has access to the full, current documentation. **Why this makes a difference:** - **Cutoff window problem solved:** The agent isn't limited to knowledge from before January 2025 - **Current API:** New frameworks, new library versions - the agent knows the current syntax - **Zero hallucinations:** Instead of "inventing" API, it reads from official documentation **Pro tip:** Index the documentation for libraries you use regularly. For me, that's: - Tailwind CSS - React (especially new hooks) - Convex - Framer Motion The agent stops being "generic AI" and becomes **an expert in your stack**. ### Hack 10: Agent Steering - Take Control on the Fly Imagine: the agent is writing a component. Halfway through, you realize you want to add another function. Wait until it finishes? Interrupt and start over? There are better options. **Agent steering** is the ability to send messages while the agent is working. You either wait until it finishes or interrupt immediately. **Where to find it:** ```text Setting location: Agent pane -> ... (three dots) -> Agent Settings -> Queue Messages Available options: 1. Send after current message -> Agent finishes the current task, then executes the new one -> Ideal for queuing tasks 2. Stop & send right away -> Immediate interruption and new task -> Use when the work direction needs to change ``` I use **"Send after current message"** **most often**. The agent finishes the current task (e.g., a component), then automatically moves to the next one on the list (e.g., tests for that component). Efficient queuing without micro-management. I use **"Stop & send right away"** **only** when: - The agent went in the wrong direction and every second is wasted - Requirements changed and the current task no longer makes sense - I need an immediate answer to a question **Caveat:** Overusing "Stop & send" causes chaos. The agent gets confused on context. Frustration rises. Use sparingly and only when you really need to change direction. **Pro workflow:** Plan a list of tasks. Send the first one normally. Queue the rest via "Send after current message." The agent works through tasks like a machine, and you can switch to something else. ## Level 3 - Pro Features These are features most users don't know exist. But professionals use them **daily**. ### Hack 11: Worktrees - Test Multiple Solutions in Parallel **Worktrees** is a feature that sounds complicated but changes everything. It lets you **test 3 different approaches simultaneously** without waiting for sequential iterations. **What are worktrees:** **Git worktrees** are isolated copies of a project on separate branches. Cursor uses this mechanism for parallel testing of multiple solutions. Each version works in its own code copy, without conflicts. When done, you pick the best solution and bring it to the main branch. **How it works step by step:** 1. **Start:** Open a chat with the agent in Cursor 2. **Choose mode:** Below the text field, you see options: **Local / Worktree / Cloud** - choose **Worktree** 3. **Choose model:** Click the model selection field 4. **Multiple models or multiplier:** Now you have two options: - **Multiple models:** Check "use multiple models" and select different models (e.g., Composer + Sonnet + GPT-4o) - **Multiplier:** Select one model and set a 2x, 3x, or 4x multiplier (e.g., 3x Claude Sonnet - you'll get 3 different versions from the same model) 5. **Automatic setup:** Cursor creates a git worktree for each version, copies the project, starts servers on separate ports 6. **Parallel work:** All versions work simultaneously on the same task 7. **Preview:** Check results on different ports (localhost:3001, 3002, 3003) 8. **Apply:** Choose the best solution -> Apply -> changes go to the main branch **Setup trick - autostart servers:** ```json // .cursor/worktrees.json - automatic setup // Place in the project root directory { "commands": [ "npm install", // Installs dependencies in worktree "npm run dev" // Starts dev server automatically ] } ``` **Practical workflow:** ```bash # Practical example of using worktrees: # 1. You have a task: "Add dark mode toggle with smooth transition" # 2. Chat with agent -> choose "Worktree" (instead of Local/Cloud) # 3. Click model selection -> check "use multiple models" # 4. Select: Composer + Sonnet + GPT-4o (or set 3x Sonnet for variety) # 5. Cursor automatically: # - Creates branch-wt-composer-xyz # - Creates branch-wt-sonnet-abc # - Creates branch-wt-gpt4o-def # - Runs npm install + npm run dev in each # 6. After 2-3 minutes you have 3 working versions: # - localhost:3001 (Composer) - minimalist toggle # - localhost:3002 (Sonnet) - animated slider with icons # - localhost:3003 (GPT-4o) - toggle with color preview # 7. Open all 3 in the browser, compare # 8. Sonnet looks best -> Review changes -> Apply # 9. Changes go to main branch, worktrees are cleaned up ``` **Why this makes a difference:** - **Design choices:** 3 models = 3 different visual approaches, pick the best - **Refactoring:** different implementation strategies, compare code quality - **Bug fixes:** see which approach is most elegant - **Zero wasted time:** Without worktrees, you'd wait 15 minutes for 3 iterations. With worktrees? 5 minutes and all versions at once. **Real use case from my workflow:** Design changes - I run 3 models simultaneously. In 5 minutes, I see 3 different visual approaches. Without worktrees, I'd wait 15 minutes for sequential iterations, and each subsequent one would be "contaminated" by feedback from the previous one. **When NOT to use worktrees:** - Simple bugfixes (overkill) - Clear requirements (one model is enough) - Limited credits (3 models = 3x cost) Worktrees are a bazooka. Use it only when you need a bazooka. ### Hack 12: Two-Model Workflow - GPT-5.2 Plans, Claude Executes **GPT-5.2 High** is the best model for planning. It creates detailed, well-thought-out plans that account for edge cases and architecture. There's just one problem: it's **slow as hell** at implementation. **Claude Sonnet** is fast at coding. Decent at planning, but not at GPT-5 level. The solution? **Use both**. ```text Optimal workflow: 1. Plan Mode -> GPT-5.2 High (plan) 2. After the plan: Switch model -> Claude 4.5 Sonnet 3. Click "Build" Result: Best plan + fast implementation ``` **How it works:** 1. Give the task in **Plan Mode** 2. Choose **GPT-5.2 High** as the model 3. Wait (yes, it takes a while) for GPT to create a detailed plan 4. **BEFORE clicking "Build"** - switch the model to **Claude 4.5 Sonnet** 5. Now click "Build" 6. Claude implements GPT's plan - quickly and efficiently **Result:** You get **GPT's planning precision** + **Claude's implementation speed**. This trick saves me **70% of the time** compared to using GPT alone for the build, or Claude alone for the plan. **Caveat:** GPT-5.2 burns through credits. Use this technique for **larger features**, not micro-tasks. For simple things, Claude alone is sufficient. This is the best example that **choosing a model doesn't have to be an "either/or" decision**. You can combine the strengths of different models in a single workflow. ### Hack 13: Structured Prompting - User Stories and Design Patterns **Problem:** Most developers write prompts ad-hoc. "Add login." "Fix this bug." "Make it prettier." The agent gets vague requirements and generates generic code. Before I discovered structured prompting, **50% of first iterations with AI were throwaway**. The code technically worked but didn't meet requirements. Because the requirements were unclear. **Solution:** Structure prompts using patterns AI models are trained on - **user stories** and **design patterns**. **User Story Format (for feature development):** ```text As a [user type] I want [goal/desire] So that [benefit/value] Acceptance Criteria: - [ ] Criterion 1 - [ ] Criterion 2 - [ ] Criterion 3 Technical Context: - Existing components: [list] - Design system: [colors, spacing] - Similar implementations: [reference files] ``` **Why this works:** - AI models are **trained on user stories** from GitHub/Jira - it's their native language - The format **forces clarity of requirements** - you can't be vague - **Acceptance criteria = built-in testing checklist** - the agent knows when it's done - **Technical context eliminates hallucinations** - the agent references existing code **Design Pattern Prompts (for architecture):** ```text I need to implement [feature] using [pattern name] pattern. Pattern Structure: - Component A: [responsibility] - Component B: [responsibility] - Communication: [how they interact] Constraints: - Must work with existing [X] - Performance requirement: [Y] - Should follow project convention: [Z] References in codebase: - Similar pattern used in: [file path] ``` **Real Examples - Bad vs Good:** **Bad prompt:** ```text Add user authentication ``` **Good prompt (user story):** ```text As a returning user I want to log in with email/password So that I can access my saved preferences Acceptance Criteria: - [ ] Login form with email + password fields - [ ] Form validation (email format, password min 8 chars) - [ ] Error handling for invalid credentials - [ ] Success: redirect to dashboard - [ ] Failure: show error message, keep user on page Technical Context: - Auth provider: Clerk (already configured) - Similar form: src/components/ContactForm.jsx - Design system: Tailwind with dark-800 backgrounds - State management: React Context in src/context/AuthContext.jsx ``` See the difference? The **bad prompt** is a guessing game. The agent has to assume 10 things. The **good prompt** is an instruction manual. The agent knows exactly what to do, how, and what "done" means. Now **80% of first-iteration code goes to production**. Not because AI got better. Because I got better at communicating with AI. **Bonus trick - Image-to-code prompts:** ```text [Attach screenshot/mockup] Recreate this design with following specifications: - Framework: React + Tailwind - Components to extract: [list] - Interactive elements: [behaviors] - Responsive breakpoints: mobile (< 768px), desktop (>= 768px) - Color palette: [specify if different from image] DO NOT: [list of things to avoid] ``` The agent gets a visual reference + technical context. Result? Pixel-perfect implementation on the first try. **When to use structured prompting:** - New features (user stories) - Refactoring (design patterns) - Design implementation (image-to-code) - Simple bugfixes (overkill) - Exploratory tasks (too much structure stifles creativity) ### Hack 14: @ Command Power Features - Context at Your Fingertips Instead of describing in words "that file with auth," I type `@ -> Files -> AuthContext.jsx` and the agent immediately has full context. Zero friction. The **@ command** is a quick way to add context to the agent. No copy-pasting code, no describing - just point. **Available options:** ```text @ command - available options: 1. Files & Folders -> Add specific files to context -> Example: @ Files -> src/components/Header.jsx 2. Docs -> Indexed library documentation -> Example: @ Docs -> Tailwind CSS 3. Terminals -> Terminal output (errors, logs) -> Useful for debugging 4. Branch (Diff with main) -> Differences between branches -> See what changed in the feature branch -> Example: "What did I add in this branch?" 5. Browser -> Content of open pages -> Design inspiration, online documentation All options can be combined for full context! ``` **Use cases:** - **Files:** "Check this file before suggesting changes" -> `@ Files -> auth.js` - **Docs:** "Use official Tailwind v4 documentation" -> `@ Docs -> Tailwind CSS` - **Branch:** "Show what changed from the main branch" -> `@ Branch (Diff with main)` - **Browser:** "Implement the design from this page" -> `@ Browser` **Pro tip:** You can add **multiple contexts at once**: "Implement authentication flow similar to auth.js, using Clerk docs, based on the changes in current branch" `@ Files -> auth.js` + `@ Docs -> Clerk` + `@ Branch` The agent has the **full picture**. It doesn't guess, doesn't hallucinate. It simply implements according to what it sees. ### Hack 15: Duplicate - Clone Well-Prepared Context For **a year**, I used Cursor and didn't know this existed. Hidden deep in the agent management view. **Use case:** You spent 30 minutes "priming" the agent. You loaded docs, rules, project context. The agent understands the codebase perfectly. Now you want to implement **3 related features** with the same foundation. Option A: Re-prime the agent 3 times. 90 minutes wasted. Option B: **Duplicate**. 3 clicks. **Where to find it:** ```text Workflow with duplicate: 1. Prime agent: load docs, rules, project context 2. When the agent understands the codebase well 3. Go to the agent management view (list of all agents) 4. Find the agent you want to clone 5. Click ... (three dots) next to that agent 6. Select "Duplicate" 7. New agent with identical context 8. Implement the next feature without repeating the priming Warning: If duplicate doesn't work: restart Cursor ``` **NOTE:** You will NOT find this option in the `...` menu directly in the agent pane during a conversation. You need to go to the **agent management view** (list of all agents in the sidebar). ### BONUS Hack 16: .cursorignore Manipulation - Dangerous but Useful **WARNING: Use ONLY in a test environment. NEVER in production.** `.cursorignore` controls what AI can see. By default, it blocks `.env` and `.env.local` files for security. And rightly so - **you don't want** API keys ending up in an AI conversation. But sometimes - in a **test repo with dummy credentials** - you want the agent to **validate the environment variable setup**. **The hack:** ```bash # .cursorignore file # WARNING: Use ONLY in a test environment # Default (safe): .env .env.local # Hack - unblocking (FOR TESTS ONLY): !.env.local # Allows AI to see environment variables # Use only with dummy/test credentials! ``` The `!` prefix negates the rule. The agent can now see `.env.local`. **Use case:** The agent can check if you have all the necessary variables, if the names are correct, if the values have the right format (of course, it doesn't see actual production values - because this is a TEST REPO). **Personal warning:** Never do this in production. Only in a test repo where keys are dummy. If in doubt - **don't do it at all**. I'm showing this hack for completeness. But use it with caution. ## Bonus Tips - Quick Pointers These didn't make the Top 15, but I use them regularly: - **Split terminals:** Hover over the terminal, click the split icon. Two terminals side by side - e.g., one for the dev server, another for git commands. - **Auto theme switching:** Editor Settings -> Auto detect color scheme. Cursor switches between light/dark theme with your system. Small detail, but convenient. - **Generate cursor rules from docs:** A slash command that extracts rules directly from project documentation. Instead of writing manually, the agent builds `.cursorrules` from the README and CONTRIBUTING. - **Design mode for prototyping:** A custom command that tells the agent to use only mock data. Perfect for prototyping UI without backend infrastructure. - **Agent review:** Automated code review on commit. The agent checks code before committing and flags potential issues. Works like a pre-commit hook with AI. Each of these tricks saves a few minutes daily. Combined - hours per week. ## Key Takeaways After going through all 16 hacks, time to summarize what truly matters: 1. **Keyboard shortcuts save hours** - `Cmd+J`, `Cmd+E`, `Shift+Tab` are the absolute basics. Stop clicking, start using the keyboard. 2. **The context window is a precious resource** - Manage MCPs, restart conversations after 60%, be aware of costs. Every unused message is wasted money. 3. **Custom commands eliminate repetition** - If you do something more than 2x, create a command. Simple as that. 4. **Documentation in Cursor = superpowers** - Index the libraries you use. The agent stops guessing and starts **knowing**. 5. **Structure prompts like professionals** - User stories and design patterns = 80% of first-iteration code to production. The biggest change in my workflow. 6. **Advanced features aren't intuitive** - Worktrees, structured prompting, duplicate chat are **hidden** but game-changing. You have to consciously seek them out. 7. **Subscription ROI grows with knowledge** - **$20/month** is little or a lot depending on how you use it. After applying these hacks, the return shows up on the first day of the month. ## Next Steps Don't try to implement everything at once. That's a path to failure. Instead: **Today (10 minutes):** - Enable **usage summary** (Settings -> Chat -> Always) - Enable **early access** (Settings -> Beta) - Enable **completion sounds** (Settings -> General) **This week (1 hour):** - Learn the shortcuts: `Cmd+J`, `Cmd+E`, `Shift+Tab` - Check which **MCPs** you have enabled, disable unnecessary ones - Start monitoring the **context window** - restart after 60% **This month (3-4 hours):** - Create your first **custom command** for a repetitive task - Index the documentation for your stack's main library - Start structuring prompts in the **user story** format **Next month (experiment):** - Try **worktrees** on your next design decision - Test the **two-model workflow** (GPT plan + Claude build) - Dive into **advanced prompt patterns** These hacks changed the way I work with AI. I started from the basics and gradually discovered deeper layers. **You can travel this path faster** - because you have this guide. Good luck. And remember - **$20 a month** is an investment, not an expense. Provided you know how to use it. ---

Want to maximize AI in your coding workflow?

I'll help you optimize your workflow with AI coding assistants, build custom automation, and train your team. From strategy through implementation to advanced techniques.

Book a free consultation
## Useful Resources - [Cursor Documentation](https://cursor.com/docs) - Official documentation - [5 Techniques for Working with Claude Code](/blog/5-technik-pracy-z-claude-code) - Complementary article on Claude Code - [Cursor Community Discord](https://discord.gg/cursor) - User community - [GitHub: claude-piv-skeleton](https://github.com/plipowczan/claude-piv-skeleton) - Workflow methodology with custom commands ## FAQ
### Is the $20/month Cursor subscription worth it for beginner developers, and how quickly does it pay for itself? Yes, the subscription pays for itself quickly thanks to access to Pro features like worktrees, unlimited Claude 3.5 Sonnet, and Composer mode. Even beginners save hours per week by avoiding manual debugging and writing boilerplate code. The $20 cost is an investment in productivity, not just a tool expense.
### What are the most important Cursor keyboard shortcuts for maintaining flow when working with AI? The absolute minimum is `Cmd+E` for switching between code and the agent, and `Shift+Tab` for quickly changing modes (Ask/Agent/Plan). It's also worth using `Cmd+J` for toggling the terminal, which eliminates taking your hands off the keyboard. These shortcuts let you treat AI as a natural extension of the editor.
### How do you effectively manage the context window limit in Cursor to avoid agent memory issues? The key is monitoring the usage indicator and resetting the chat after exceeding 60% capacity, which prevents model hallucinations. You should also disable unused MCPs in settings, since each active add-on consumes precious tokens. Using `@Files` precisely instead of entire folders also saves context.
### What is working with worktrees in Cursor and when is it best applied? Worktrees enable parallel generation and testing of multiple code variants (e.g., by different AI models) on isolated branches. Cursor automatically sets up environments for each version, letting you compare, say, three UI approaches in minutes. It's the ideal tool for experimentation and architectural decision-making.
### Why is it worth adding custom documentation to Cursor's index instead of relying on the model's general knowledge? AI models often have outdated knowledge about the latest library versions (e.g., Tailwind v4), leading to incorrect code generation. Adding documentation URLs in "Docs" ensures the agent uses a "single source of truth" and doesn't invent non-existent functions. This eliminates the "cutoff date" problem and improves generated code quality.
--- # 5 Techniques That Will Transform How You Work with Claude Code Source: https://pawel.lipowczan.pl/en/blog/5-techniques-working-with-claude-code Published: 2026-01-09 There's a very high probability that you're leaving most of your AI coding assistant's potential on the table. When I started working with Claude Code to build this portfolio, I was doing exactly the same thing - typing simple prompts, getting code back, sometimes it worked, sometimes it didn't. Reactive prompting. No system. Then I discovered that the best AI engineers work completely differently. They have a **system**. They use methodologies that make their coding agents more powerful with every iteration. In this article, I'll show you 5 concrete techniques that completely change how you work with Claude Code. These aren't theoretical concepts - they're practical methods used by teams building production applications with AI. And the best part is that all these techniques are already packaged into a ready-to-use framework you can deploy in your project today. ## PRD-First Development: Your Project's North Star Most developers just dive into code. They open Claude Code and start: "Add a login button," "Add form validation," "Fix this bug." Each iteration is disconnected from the previous one. There's no big-picture vision. A **PRD (Product Requirement Document)** in the context of working with AI is something much simpler than a 50-page corporate document. It's simply a **markdown file with the full project scope**. A single file that becomes the north star for every feature you build. ### Benefits of PRD-First Development Previously, I used an assistant that generated PRDs and rules for the AI agent, but it required working with two different tools. I had to switch between contexts, sync information manually. It was frustrating. Now? Everything in one place. The PRD lives in my repository. Claude Code reads it at the start of every session. And suddenly everything makes sense. ### What It Looks Like in Practice For new projects (greenfield development), the PRD contains: - **Target Users** - who you're building for - **Mission** - what the product should do - **In Scope / Out of Scope** - what's in the MVP, what's for later - **Architecture** - high-level tech stack and structure For existing projects (brownfield), the PRD documents: - **What we already have** - current system state - **What we're building next** - upcoming features in the pipeline - **Long-term vision** - where we're heading Minimal PRD structure example: ```markdown # PRD: Habit Tracker Application ## Target Users People who want to build better habits through consistent tracking ## Mission A simple, elegant app for tracking habits with progress visualization ## In Scope (MVP) - Creating habits with name and frequency - Marking daily habit completion - Calendar showing history (streak tracking) - Local data storage ## Out of Scope (v1) - Sharing with other users - Advanced statistics - Push notifications - Integrations with other apps ## Architecture - Frontend: React + TypeScript - State: React Context - Storage: localStorage - Deploy: Vercel ``` ### The Magic Question When you have a PRD, you can start every session with a question that changes everything: > "Based on PRD, what should we build next?" Claude Code reads the PRD, understands where you are, what you've already built, and suggests the next logical step. You don't have to remember. You don't have to explain context from scratch. **The PRD remembers for you**. ### Key Benefits - **Single source of truth** - one project definition for the entire team (and for AI) - **Natural decomposition** - easy to extract the next features to implement - **Context for the agent** - Claude Code always knows what you're working on - **Bird's eye view** - every feature connects to the bigger vision The PRD is the foundation. Without it, you're building a house on sand. With it - you have a solid foundation, and every line of code has purpose and direction. ## Rule Modularity: Lighter Context, Smarter Agent I've seen this dozens of times. A developer creates a `CLAUDE.md` or `agents.md`, dumps every possible rule, guideline, and convention in there. After two months, the file is 1,500 lines long. All of it gets loaded **at the start of every conversation**. The problem? **You're overwhelming the LLM with irrelevant context.** ### The Problem with Long Global Rules When you're working on the frontend, you don't need to know API design patterns. When you're working on the database, you don't need React component naming conventions in context. But if everything is in one global rules file? Claude Code loads it all. Every time. That's wasting the **context window** - something many developers seriously underestimate, but which is absolutely critical for agent output quality. ### The Solution: Modular Architecture Instead of one monstrous file, split your rules into two categories: **1. Global rules (CLAUDE.md)** - keep it as light as possible, ~200 lines - Project tech stack - Folder structure - Commands to run (npm run dev, npm test) - Testing strategy (philosophy, not details) - Logging standards **2. Reference folder** - detailed context loaded only when needed - `reference/api-design.md` - REST API patterns, error handling, loaded only when working on APIs - `reference/frontend-components.md` - component patterns, styling guidelines, only for UI work - `reference/database-patterns.md` - schema design, migrations, query optimization, only for DB work - `reference/testing-patterns.md` - detailed test examples, setup, mocking ### How to Set It Up In your main `CLAUDE.md`, add a reference section: ```markdown # CLAUDE.md ## Tech Stack - React 19 + TypeScript - Tailwind CSS 3 - Vite 7 ## Project Structure src/ ├── components/ ├── pages/ ├── utils/ └── data/ ## Reference Documentation When working on specific areas, consult these documents: - **API endpoints**: `.claude/reference/api-design.md` - **Frontend components**: `.claude/reference/frontend-components.md` - **Database operations**: `.claude/reference/database-patterns.md` - **Testing**: `.claude/reference/testing-patterns.md` ``` Claude Code is smart enough to understand: "OK, I'm working on an API endpoint now, I should read api-design.md." And it does so automatically. ### Benefits of Rule Modularity Having everything in one place (like in claude-piv-skeleton, which I'll tell you about shortly) vs. scattered across different tools is an **enormous difference**. Everything is in the repo. Everything is versioned. Everything evolves alongside the project. And most importantly? **You're protecting the context window** for things that actually matter during the implementation of a specific feature. ## Commandification: Stop Repeating the Same Prompts If you're sending the same prompt to your coding agent more than twice, that's a screaming signal: **"Turn me into a command!"** ### What Are Commands Commands are simply **markdown files that define a workflow**. You load them as context, and Claude Code executes the defined process step by step. They're like macros for prompts. This doesn't require any new tools. It's just a better way of organizing your work. ### What's Worth Commandifying Practically everything you do regularly: **Core workflow:** - `/prime` - Load project context at the start of a session - `/plan-feature` - Create a structured implementation plan for a feature - `/execute` - Implement according to the plan - `/validate` - Run tests and verify functionality **Git operations:** - `/commit` - Create a meaningful commit message based on changes - `/review-pr` - Analyze a pull request before merging **Maintenance:** - `/update-docs` - Update documentation after code changes - `/refactor` - Refactor according to project best practices ### Example: The /prime Command ```markdown # Command: /prime Load codebase context to prepare for feature development work. ## Objective Ensure Claude Code understands current project state before starting any development. ## Steps 1. Read PRD from `docs/PRD.md` to understand project scope and vision 2. Read architecture documentation from `docs/ARCHITECTURE.md` 3. Scan recent commits: `git log -10 --oneline` to see latest changes 4. List current todos from `docs/TODO.md` to understand priorities 5. Check git status to see any uncommitted work 6. Confirm context successfully loaded with summary of project state ## Expected Output Brief summary including: - Project name and current phase - Last 3 features implemented - Current priorities from TODO - Any blockers or issues noted ``` Instead of typing out every time: "Read the PRD, then the architecture, then check what changed recently..." - you just call `/prime` and you're done. ### Key Observation Since the assistant is essentially just a prompt, you can use that prompt as a command. That's the beautiful property of markdown commands - they're portable, human-readable, and you can iteratively improve them. ### Typical Workflow with Commands Here's what my daily work cycle with Claude Code looks like: ```text /prime -> "Based on PRD, what should we build next?" -> /plan-feature "Add user authentication" -> [Context Reset - new conversation] -> /execute plan-auth.md -> /validate -> /commit ``` ### Benefits of Commandification - **Consistency** - the same process every time, zero missed steps - **Zero forgetting** - you don't have to remember the sequence, the command remembers - **Onboarding** - a new developer on the team gets ready-made commands and is productive immediately - **Continuous improvement** - commands evolve and get better over time You literally save **thousands of keystrokes** per year. And more importantly - you save mental energy for things that actually require thinking. ## Context Reset: The Most Important Step You're Not Taking This sounds counterintuitive: **Always reset the conversation between planning and execution.** Most developers do this wrong. They plan a feature in a long conversation with Claude Code - reading files, discussing architecture, exploring different approaches. And then, in the same conversation, they immediately start implementing. The problem? **The context window is cluttered** with the entire exploration process. ### Benefits of Context Reset During planning, you load TONS of context: - Reading many files from different parts of the project - Exploring existing architecture - Discussing various approaches - Reviewing similar implementations in the codebase But during execution, you want: - **Maximum reasoning space** for the LLM - **Room for self-validation** - the agent should be able to check its own solutions - **A clean mental model** - only what's needed to implement this specific feature ### The Right Workflow ```text [Planning Session] -> /prime - Load codebase context -> Conversation: "Based on PRD, let's plan authentication feature" -> /plan-feature - Output: structured plan as a markdown document -> [NEW CONVERSATION] <- This is the key! -> /execute plan-auth.md - The ONLY context is the plan -> Implementation with full reasoning space ``` ### What the Plan Document Contains The plan must be self-contained - containing EVERYTHING needed for implementation: ```markdown # Plan: User Authentication Feature ## Feature Description Implement JWT-based authentication with email/password login. ## User Story As a user, I want to securely log in to access personalized content. ## Context to Reference - `src/utils/api.js` - API utility functions - `src/context/AuthContext.jsx` - existing auth context (modify) - `docs/reference/api-design.md` - API patterns ## Technical Approach 1. Backend: Create /api/auth/login and /api/auth/register endpoints 2. Frontend: Login form component with validation 3. State: Store JWT token in AuthContext 4. Protected routes: Add authentication middleware ## Task-by-Task Breakdown ### Task 1: Backend Auth Endpoints - Create `src/api/auth.js` with login and register functions - Implement JWT token generation - Add password hashing with bcrypt - Error handling for invalid credentials ### Task 2: Login Form Component - Create `src/components/LoginForm.jsx` - Form validation using Formik - Connect to auth API - Handle loading and error states ### Task 3: Auth Context Updates - Modify `src/context/AuthContext.jsx` - Add login/logout/register methods - Persist token to localStorage - Auto-refresh token logic ### Task 4: Protected Routes - Create ProtectedRoute component - Redirect to /login if not authenticated - Update router configuration ## Testing Requirements - Unit tests for auth API functions - Integration tests for login flow - E2E test: complete registration and login - Error handling tests: invalid credentials, expired token ## Success Criteria - User can register with email/password - User can login and access protected pages - Token persists across page refreshes - Logout clears token and redirects to home ``` ### Benefits of a Self-Contained Plan The LLM has **maximum token count** for reasoning during the critical coding phase. There's no contamination from exploratory context. And most importantly - **it forces you to create complete, self-contained plans**. It's discipline. But discipline that makes your AI coding sessions incomparably more effective. I've tried it both ways. The difference is enormous. Context reset is the technique I took the longest to learn, but which gave the biggest boost in output quality. ## System Evolution: Every Bug Is a Lesson for the Agent This is the most important technique of all. And the most commonly overlooked. For me, it was unknown until recently. Perhaps for you too, or you don't realize its significance. The traditional approach looks like this: 1. The AI agent makes a mistake 2. You fix it manually 3. You move on 4. **The same mistake repeats a week later** The evolutionary approach: 1. The AI agent makes a mistake 2. You analyze: **What in the system allowed this mistake?** 3. You update rules/commands/process 4. **This class of bugs is eliminated forever** ### The Mindset Shift Don't fix the bug. **Fix the system that allowed the bug.** Your AI agent is not a static tool. It's an **evolving system** that can become more powerful with every iteration. But only if you actively improve it. ### Examples of System Evolution **Scenario 1: Wrong import styles** ```text Bug: Agent uses require() instead of import in an ES6 project Analysis: No clear rule about the module system Fix: Add to CLAUDE.md: "Always use ES6 import/export syntax. Never use require() or module.exports." Result: Never again a problem with import styles ``` **Scenario 2: Forgets to run tests** ```text Bug: Agent implements a feature without tests Analysis: No "testing" step in the workflow Fix: Update the /execute command template: ## Testing Phase 1. Write tests first (TDD when appropriate) 2. Run full test suite: npm test 3. Ensure all tests pass before completion 4. Add test coverage report to plan Result: Tests become an automatic part of the workflow ``` **Scenario 3: Doesn't understand the auth flow** ```text Bug: Agent implements incorrect authentication Analysis: No documentation of the authentication flow Fix: 1. Create reference/authentication.md with a flow diagram 2. Add to CLAUDE.md: "When working on authentication, read reference/authentication.md" Result: Auth implementations are consistently correct ``` ### The Reflection Workflow After finishing each feature, instead of immediately rushing to the next one: ```markdown "Hey Claude, I noticed that XYZ wasn't working correctly and I had to fix it. Let's analyze: 1. Read the commands we used 2. Read the current rules 3. Identify what we can improve so this doesn't repeat 4. Suggest specific changes to rules/commands" ``` Claude Code analyzes the session, finds gaps in the process, and proposes fixes. Sometimes it's a new rule. Sometimes an extra step in a command. Sometimes a new reference document. ### Benefits of System Evolution - **Your agent gets smarter over time** - reliability grows with every iteration - **You build institutional knowledge** - the whole team benefits from learning from mistakes - **Transformation from reactive to proactive** - instead of fighting fires, you prevent them - **Compound effect** - after 3 months, you have a system that makes fewer mistakes than a junior developer This is more than a technique. It's a **mindset**. Treat your AI agent's system like production code that requires continuous improvement. From my own experience, the biggest mistake is ignoring patterns in agent errors. The first time is an accident. The second time is a signal. The third time is your fault for not fixing the system. ## Claude PIV Skeleton: Everything in One Place Now the best part. All these techniques I've been talking about - PRD-first development, rule modularity, workflow commandification, context reset, system evolution - are already implemented in a ready-to-use framework. It's called **Claude PIV Skeleton** and is available on GitHub: [https://github.com/plipowczan/claude-piv-skeleton](https://github.com/plipowczan/claude-piv-skeleton) (fork from [galando](https://github.com/galando/claude-piv-skeleton)) ### The Problem It Solves Previously, I used an assistant that generated PRDs and rules for the AI agent, but it required using two different tools. I had to copy output between applications, sync manually, lose context. Having everything in one place is an **enormous advantage**. And since the assistant is essentially just a prompt, you can use that prompt as a command - and that's exactly what PIV Skeleton does. ### What Is the PIV Methodology PIV stands for **Prime-Implement-Validate** - a methodology created by [Cole Medin](https://github.com/coleam00) specifically for development with AI assistants: - **Prime**: Load and understand the codebase context - **Implement**: Plan features and execute implementation - **Validate**: Automatically test and verify This isn't an abstraction. It's a concrete workflow that Cole proved in production projects like [woningscoutje.nl](https://woningscoutje.nl). ### What You Get in claude-piv-skeleton The repository implements all 5 techniques as a ready-to-use framework: **Universal methodology** - works with any tech stack (Spring Boot, Node.js, React, Python FastAPI, etc.) **Modular rules system** - path-based rule loading depending on what you're working on **Pre-built commands** - a complete set of commands for the entire PIV workflow **Technology templates** - ready-made configurations for popular stacks ### Repository Structure ```text .claude/ ├── CLAUDE.md # Lightweight global rules ├── PIV-METHODOLOGY.md # Full methodology documentation ├── commands/ # All workflows as commands │ ├── piv_loop/ # Core PIV workflow │ │ ├── prime.md # Prime phase command │ │ ├── plan-feature.md # Planning command │ │ └── execute.md # Execution command │ ├── validation/ # Validation & testing │ │ ├── validate.md # Full validation pipeline │ │ ├── code-review.md # Technical review │ │ └── system-review.md # Process improvement │ └── bug_fix/ # Bug fix workflow │ ├── rca.md # Root cause analysis │ └── implement-fix.md # Fix implementation ├── rules/ # Modular rules by technology │ ├── 00-general.md # Universal principles │ ├── 10-git.md # Git workflow │ ├── 20-testing.md # Testing philosophy │ └── backend/ # Backend-specific rules └── reference/ # Best practices loaded on-demand └── patterns/ # Design patterns reference ``` ### The Workflow It Enables Complete development cycle with PIV Skeleton: ```bash # 1. Prime workspace "Run /piv_loop:prime to load the project context" # 2. Plan feature "Use /piv_loop:plan-feature to create a plan for adding user authentication" # 3. Execute (automatic context reset!) "Use /piv_loop:execute to implement the plan" # 4. Validation runs automatically # No manual step needed - testing happens in the workflow # 5. Bug fix with system evolution built in "Run /bug_fix:rca for issue #123" "Use /bug_fix:implement-fix to implement the fix" ``` ### Benefits of PIV Skeleton - **You don't have to build from scratch** - ready structure, tested in production - **Battle-tested patterns** - workflow developed by top AI engineers - **Community-driven** - contributions from many developers, continuous improvements - **Extensible** - easy to adapt to your tech stack and needs PIV Skeleton isn't just code. It's a **system of thinking** about working with AI agents, packaged into a reusable framework. ## How to Start: First Steps Great, you know the 5 techniques and you know there's a ready framework. But how do you actually put it all into practice? ### Path 1: Use PIV Skeleton (Recommended) If you're starting a new project or can migrate an existing one: ```bash # Clone the repository git clone https://github.com/plipowczan/claude-piv-skeleton.git my-project cd my-project # Remove git history to start from a clean slate rm -rf .git git init # Install your tech stack # (Follow the technology-specific guides in the technologies/ directory) # Start your first feature # Open Claude Code and: ``` 1. `"Run /piv_loop:prime to load project context"` 2. `"Based on PRD, what should we build first?"` 3. `"Use /piv_loop:plan-feature to plan it"` 4. `"Use /piv_loop:execute to implement"` And that's it. You have a working system with your first feature in production. ### Path 2: Implement Incrementally If you have an existing project and don't want to move everything at once, adopt the techniques step by step: **Week 1: Create a PRD** - Document the current state of the project - Define the next features to build - Make the PRD your north star **Week 2: Create a Prime Command** - What context should always be loaded? - Create a `/prime` command in `.claude/commands/` - Use it at the start of every session **Week 3: Modularize Rules** - Split your CLAUDE.md into global + reference - Move task-specific rules to `reference/` - Add a reference section to global rules **Week 4: Add Feature Workflow** - Create a `/plan-feature` command - Practice context reset between planning and execution - Create an `/execute` command **Week 5: System Evolution** - After every bug, do a reflection - Update rules/commands based on findings - Track improvements in CHANGELOG ### Key Success Factors 1. **Start small** - Don't try to implement all 5 techniques at once. Begin with the PRD, then add the prime command, etc. 2. **Document as you go** - Write down what works and what doesn't. Your notes will become part of the system evolution. 3. **Iterate on commands** - Your workflows will improve. That's normal. After a month, your `/prime` will be better than at the start. 4. **Be consistent** - Use the system every time. Don't fall back into old habits of "quick fixing" without process. 5. **Share with the team** - If you work in a team, make sure everyone uses the same commands and processes. You multiply the benefits. ### Your First Feature with PIV A concrete example - let's say you're building a habit tracker and want to add streak tracking: ```text Session 1 - Planning: -> /prime -> "Based on PRD, let's plan streak tracking feature" -> /plan-feature "Streak tracking - show consecutive days" -> Plan saved: .claude/agents/plans/streak-tracking.md [Restart conversation] Session 2 - Execution: -> /execute .claude/agents/plans/streak-tracking.md -> [Implementation happens with tests] -> /validate -> All tests pass Session 3 - Commit: -> /commit -> "feat: Add streak tracking with visual indicators" ``` In an hour, you have a feature in production. With tests. With a proper commit message. With everything. ## Key Takeaways 1. **PRD-first development** ensures consistency and direction for all iterations with the AI agent. It's the north star that makes every feature meaningful in the context of the whole. 2. **Rule modularization** protects the context window and loads only needed knowledge. Stop wasting tokens on irrelevant context - load what matters, when it matters. 3. **Workflow commandification** saves thousands of keystrokes and ensures consistency. If you do something more than twice, it should be a command. 4. **Context reset** between planning and execution gives the agent maximum reasoning space. Counterintuitive, but one of the most impactful techniques. 5. **System evolution** turns every bug into a lesson that makes the agent smarter. Don't fix the bug - fix the system that allowed it. 6. **PIV Skeleton** offers a ready implementation of all techniques in one place. You don't have to build from scratch - you can start today. 7. Most importantly: **a systematic approach** vs. reactive prompting is the difference between using 20% and 80% of Claude Code's potential. ### My Experience When I started building this portfolio with AI assistance, I had no system. Simple prompts, ad-hoc fixes, zero processes. Then I discovered these techniques. And everything changed. Now my workflow with Claude Code is predictable. Effective. And most importantly - **the agent gets better with every session**, instead of making the same mistakes over and over. This is a transformation you can have on your team. It requires a mindset shift from "AI is a faster Google" to "AI is an evolving development partner." But if you make that shift? The difference will be enormous. ## FAQ
### How do I start with PRD-first development if I've never created project documents before? Start with a minimal PRD with four sections: Target Users (who it's for), Mission (what it does), In Scope (MVP features), and Out of Scope (what's for later). You don't need a 50-page document - a simple markdown with key decisions is enough. A PRD for a small project can literally be 20-30 lines and already delivers enormous value as a single source of truth for AI.
### Exactly how many rules should I have in the main CLAUDE.md file to avoid overwhelming the LLM context? A maximum of 200 lines in the main CLAUDE.md - tech stack, project structure, basic commands, and links to reference docs. Move all detailed patterns (API design, component patterns, testing) to separate files in a reference folder. Claude Code will automatically load them only when you're working on that area, saving precious context window space.
### Does workflow commandification only work with Claude Code, or can I use the same commands with ChatGPT or other LLMs? Commands are plain markdown files with workflow instructions, so they work with any LLM (ChatGPT, Claude, Cursor, Windsurf). The only difference is how you load them - in Claude Code it's slash commands, in ChatGPT you copy the contents as a prompt. The methodology and command structure itself is universal and portable across tools.
### Why do I need to reset context between planning and execution instead of doing everything in one session? During planning, you load TONS of exploratory context (reading many files, discussing various approaches), which clutters the LLM's context window. A reset gives the agent a clean slate with maximum reasoning and self-validation space during implementation. It's counterintuitive, but empirically delivers significantly better results - the agent has room for quality checks instead of fighting with an overloaded context.
### What exactly makes claude-piv-skeleton different from just using Claude Code without any system? PIV Skeleton is a ready framework with predefined commands (/prime, /plan-feature, /execute, /validate), folder structure (.claude/commands, .claude/agents), document templates (PRD, rules), and a system evolution process. Instead of inventing a workflow from scratch, you get a proven system used by teams in production - you simply fork the repo and have ready-made best practices. It's like the difference between writing your own framework and using Next.js.
### How long does it realistically take to implement these 5 techniques in an existing project that's been running for several months? Start small - day 1: create a minimal PRD (1-2h), day 2: lightweight CLAUDE.md + one /prime command (1-2h), week 1: add /plan-feature and /execute (2-3h total). Don't implement everything at once. After 2 weeks of working with the system, you'll see natural places to add more commands and rules. Brownfield projects require about 5-8 hours total setup, but you see the ROI after your very first session with the new workflow.
---

Want to implement AI in your team?

I'll help you build a system for working with AI agents that boosts your team's productivity. From strategy through implementation to training.

Book a free consultation
## Useful Resources - **[claude-piv-skeleton](https://github.com/plipowczan/claude-piv-skeleton)** - Ready framework implementing the PIV methodology - **[habit-tracker](https://github.com/coleam00/habit-tracker)** - Original PIV demo project by Cole Medin - **[context-engineering-intro](https://github.com/coleam00/context-engineering-intro)** - Introduction to context engineering by Cole Medin - **[Claude Code Documentation](https://claude.ai/code)** - Official Claude Code documentation --- # 2026: The year AI moved from labs to factory floors Source: https://pawel.lipowczan.pl/en/blog/ai-trends-2026-from-experiments-to-operationalization Published: 2026-01-01 We're sitting on the first day of 2026. If you're a technology leader, a decision maker at a company, or simply someone trying to keep up with the AI revolution, you're probably feeling a mix of excitement and uncertainty. And rightfully so. Because 2026 is the moment when AI stops being a "fascinating technology of the future" and becomes an operational foundation - a tool that we either integrate into our business processes or get left behind. After years of experiments, pilots, and "wow effect" presentations, the time for verification has come. Gartner analysts call it the "superintelligence cycle," Forrester talks about "the reckoning," and Deloitte about "the infrastructure reckoning." I simply call it: **the end of AI tourism and the beginning of real work**. ## From chatbots to agents: AI that acts instead of just answering The biggest shift I'm observing in 2026 is the move from generative models as "smart assistants" toward **Agentic AI** - systems that don't just answer questions but autonomously plan, make decisions, and take actions. ### What does this actually mean in practice? Imagine you ask AI: "Conduct a GDPR compliance audit for a new marketing process." Earlier models (GPT-4, early versions of Claude) would give you a list of steps to follow. **Agents in 2026 execute those steps on their own**: 1. A "researcher" agent analyzes the process documentation 2. A "lawyer" agent checks compliance with GDPR regulations 3. A "codifier" agent generates checklists and reports 4. A "critic" agent verifies conclusions and escalates doubts to a human This is no longer science fiction. IDC forecasts that by 2029, agentic systems will account for nearly 50% of all AI spending. By the end of 2026, 80% of enterprise applications will contain embedded AI agents. ### Multi-agent orchestration: AI teams in action The key architectural shift is the move from single, monolithic models to **multi-agent systems**. Instead of one "super-brain," we have a team of specialized agents that collaborate through protocols like Model Context Protocol (MCP) or Agent-to-Agent (A2A). In practice I've already seen deployments in: - **Customer service**: Agent systems in banking provide 24/7 support, automatically coordinating actions between departments - **Marketing**: AI generates campaign briefs, segments audiences, tests variants, and optimizes budgets without human involvement - **Software development**: Agents automate tests, file bugs, generate documentation, and even propose code fixes But heads up - **Agentic AI's success depends on "bounded autonomy."** Agents must operate within strictly defined security frameworks, with escalation mechanisms to humans in case of anomalies. This is the answer to the hallucination problems of earlier models. ## Reasoning models: AI that "thinks" before answering If Agentic AI is a revolution in action, then **Reasoning Models** are a revolution in thinking. In 2026, AI no longer generates answers instantly - instead, it spends additional compute time on "deliberation." ### System 2: slow, analytical AI thinking The term "System 2" comes from cognitive psychology and refers to thought processes that are slow, analytical, and logical (as opposed to fast, intuitive System 1). New-generation models - successors to OpenAI o1/o3, Google Gemini in advanced versions - use **inference-time compute**: they run internal simulations, verify hypotheses, and plan solution steps before generating the final answer. The results? AI achieves expert level in: - Solving mathematical problems at 93%+ accuracy (GPQA Diamond) - Multi-step tasks requiring logical reasoning - Verifying assumptions and detecting errors in argumentation Microsoft describes this as a shift from "AI as a tool" to "AI as a partner" that not only executes commands but also contributes substantive input. ### The cost of intelligence But there's a catch. Gemini 3 with reasoning enabled consumed 160 million tokens where without reasoning 7.4 million sufficed. This illustrates the fundamental **trade-off between speed and intelligence** that organizations must actively manage in 2026 through "reasoning budgets" tailored to specific tasks. ## Model specialization: the end of the "one model for everything" era One of the most important trends of 2026 is the mass shift from large, general-purpose language models (LLMs) to **Domain-Specific Language Models (DSLMs)**. ### Why does specialization win? In regulated industries - healthcare, finance, law - **accuracy matters more than universality**. DSLMs offer: - **Higher precision**: Med-PaLM achieves 95% accuracy in medical diagnostics, FinGPT reduces fraud detection by 30%, JurisGPT analyzes contracts 25-30% more accurately than general-purpose LLMs - **Lower operational costs**: Fewer parameters means inference cost reduction of up to 45% - **Built-in regulatory compliance**: Models trained on dedicated datasets include compliance mechanisms "out of the box" Gartner forecasts that by the end of 2026, over 50% of GenAI models used by enterprises will be domain-specific. In regulated sectors, that figure reaches 80-90%. ### Hybrid architecture as the standard In practice I don't see a total replacement of LLMs by DSLMs. Instead I observe a **hybrid architecture**: general-purpose models for broad tasks + domain modules for specialist functions. Cloud providers (AWS, Azure, Google Cloud) already offer dedicated platforms: Healthcare AI, Financial Services AI, Manufacturing AI - each pre-trained on curated datasets with built-in compliance frameworks. ## Infrastructure: from cloud to edge If models are AI's brain, then infrastructure is its body. And in 2026, that body is undergoing a dramatic transformation. ### The economics of inference vs. training The key shift: whereas previously most compute was consumed by **model training**, the dominant workload is now **inference**. Deloitte predicts that inference will account for two-thirds of total AI compute demand. This drives a boom in: - **ASIC chips** (Application-Specific Integrated Circuits): AWS Trainium/Inferentia, Google TPU v6, Microsoft Maia - optimized for specific model architectures, offering better performance-to-energy ratios - **Three-layer hybrid architecture**: 1. **Public cloud**: flexibility for training and experiments 2. **On-premises**: stability for critical inference and data sovereignty compliance 3. **Edge**: ultra-low latency, privacy, resilience against central service outages ### Edge AI and TinyML: intelligence everywhere Edge computing and technologies like TinyML (Tiny Machine Learning) are redefining data processing in 2026. ML models running on **microcontrollers and IoT devices** enable: - **Real-time analysis** without sending data to the cloud - **Low energy consumption** (Intel Loihi 2 neuromorphic architectures consume orders of magnitude less energy) - **Privacy protection** (medical and financial data stays local) Practical applications are already live: - Smart agriculture: edge-deployed models monitor crops in real time - Predictive maintenance: anomaly detection directly on sensors - Wearables: health analysis without constantly sending data to the cloud ### AI PCs and Small Language Models In 2026, Gartner predicts that 55% of all new computers will be **AI PCs** equipped with dedicated NPU (Neural Processing Unit) chips exceeding 40-50 TOPS performance. In parallel, the mobile market is experiencing a renaissance thanks to **Small Language Models (SLMs)** - Google Gemini Nano, Apple Intelligence - ranging from 1 to 7 billion parameters, optimized for mobile processors. Over half of new smartphones in 2026 have native GenAI support, enabling RAG (Retrieval-Augmented Generation) functions directly on the phone. This creates a new quality of "personal AI" that knows the user's context but doesn't share it with corporations. ## Physical AI: from demonstration robots to production 2026 is the moment when **Physical AI** - artificial intelligence with a body - enters factory floors and warehouses at commercial scale. ### Humanoid robots: Tesla Optimus, Figure AI, Digit - **Tesla Optimus**: Elon Musk targets 2026 as the moment for serial production launch and availability to external customers. Robots take over simple, repetitive, and dangerous tasks - **Figure AI + BMW**: The partnership reaches maturity - Figure 02 robots work autonomously on BMW assembly lines, performing manipulation tasks requiring precision - **Agility Robotics (Digit)**: The robot known for working in Amazon and GXO logistics centers achieves operational scalability through autonomous docking and WMS system integration The key is the "universal robot brain" - an AI model that allows a machine to learn new tasks through observation rather than imperative programming. ### Software-Defined Factory In manufacturing, physical automation is integrating with digital intelligence. The **Software-Defined Factory** concept assumes that production line functionality is defined by software. IDC forecasts that by 2029, 30% of factories will be managed by open automation platforms. AI is no longer just an add-on for predictive maintenance - it's becoming an **autonomous system managing production scheduling** - over 40% of manufacturers will modernize their planning systems with AI modules that dynamically respond to supply chain disruptions. ## EU AI Act: August 2026 - zero hour for compliance For companies operating in Europe, the most important calendar date is **August 2, 2026** - the deadline for full implementation of EU AI Act provisions concerning high-risk AI systems. ### What does this mean in practice? From that day, companies must have implemented: 1. **AI risk management systems** 2. **Human oversight mechanisms** (Human-in-the-loop) 3. **Training data quality assurance** 4. **Complete technical documentation and system logs** 5. **Incident reporting procedures** Non-compliance? Penalties reach **35 million euros or 7% of global turnover** - putting AI compliance on par with GDPR as a management priority. ### Polish implementation The draft law implementing the AI Act provides for establishing a Commission for the Development and Safety of Artificial Intelligence and launching the first regulatory sandbox by August 2026. This is an opportunity for Polish companies to safely test AI solutions under controlled conditions. ## Cybersecurity: from reactive to preventive The 2026 threat landscape is dominated by AI-assisted attacks. **Audio and video deepfakes** are being used to bypass biometrics and conduct advanced phishing (fake video conferences with executives). ### AI security platforms Gartner forecasts that by 2028, over half of companies will deploy **AI Security Platforms** that: - Centralize visibility of all AI systems in an organization - Enforce AI usage policies (AI Usage Control) - Protect against AI-specific threats: prompt injection, data leakage, agent manipulation - Monitor activities in real time (Runtime Monitoring) ### Digital provenance and the fight against disinformation A key defensive technology is **Digital Provenance** and C2PA standards, which allow cryptographic attestation of the authenticity and source of multimedia content. Solutions like Google SynthID, Adobe C2PA, and Microsoft GUID enable identification of AI-generated content. ### Quantum threat and post-quantum cryptography "Harvest Now, Decrypt Later" scenarios are becoming real - data encrypted today may be decrypted by quantum computers in the future. Poland is implementing **Post-Quantum Cryptography (PQC)** projects based on Kyber, Dilithium, Falcon, and SPHINCS+ algorithms. ## No-Code/Low-Code: citizen developers take the wheel In 2026, 70-75% of new enterprise applications will contain **no-code or low-code** components, compared to 25% in 2023. That's a 3x increase in five years. ### Why is this happening? - **Speed**: development time reduced by 50-70% - **Cost**: reduction of 40-60% - **Democratization**: enabling solution creation by "citizen developers" - business employees without coding skills The key 2026 shift is **AI-assisted development**. Platforms like Microsoft Power Platform allow generating application logic, workflows, and data connections from natural language prompts. ### Challenge: governance at scale The biggest challenge: maintaining quality, security, and compliance without developer gatekeeping. Leading organizations are implementing **"governed citizen development"** - oversight frameworks enabling rapid innovation while maintaining security, compliance, and architectural consistency standards. ## AI economics: ROI verification and new pricing models 2026 will bring a "reckoning" to the market. After years of enthusiasm, **investors and boards will demand hard evidence of return on investment**. ### The end of the "AI tourism" era Forrester predicts that 25% of AI budgets will be pushed to 2027 due to delays in value verification. CIOs will be forced to rescue AI projects that failed due to lack of technical competence. In 2026, only solutions delivering **measurable efficiency improvements, cost reduction, or revenue growth** will matter. IDC forecasts that 70% of G2000 CEOs will focus AI ROI on revenue growth, not just headcount reduction. ### SaaS pricing model shift The traditional seat-based pricing model is becoming obsolete in a world where work is performed by **digital agents, not people logging into systems**. IDC predicts that by 2028, 70% of software vendors will need to rebuild their pricing models, shifting to: - **Outcome-based**: paying for business results - **Consumption-based**: paying for resource usage - **Agent-based**: paying per number of active AI agents ## Labor market: role redefinition, not elimination AI in 2026 doesn't eliminate professions wholesale - it redefines job roles and requires new competencies. ### New roles in the AI era | New job role | Key competencies | Application | | ------------------ | ---------------------------------------- | ------------------------ | | AI Product Owner | AI lifecycle management, compliance | AI model deployment | | AI Risk Officer | Risk management, ethics, audit | Incident monitoring | | AI Orchestrator | Agent coordination, system integration | Process automation | | Prompt Engineer | Optimizing model interactions | All departments | | Data Curator | Training data management | AI/ML teams | ### Most threatened vs. resilient positions **Threatened**: routine cognitive tasks - data entry, basic coding, administration, first-level customer service. **Resilient**: work requiring complex judgment, empathy, creativity, and deep domain expertise. ### Reskilling as strategy Amazon is implementing its Career Choice program, and the World Economic Forum promotes Human-Machine Collaboration initiatives. Companies that succeed in 2026 treat transformation as **intentional reskilling, not reactive headcount elimination**. ## Poland 2026: where we are and where we're heading ### Digital strategy and AI factories The Ministry of Digital Affairs is finalizing plans for 2026-2027, key projects: - **mObywatel**: the app as a central hub for government services - **e-Doręczenia**: full deployment of digital official correspondence - **AI Factories**: launching computing centers in Poznan and Krakow supporting Polish researchers and SMEs - **AI Gigafactory**: a cluster project of leading centers (Poznan, Krakow, Wroclaw, Warsaw, Gdansk) - 5 billion PLN investment, 2 billion PLN from public funds ### Competency gap and adoption Despite progress, Poland is catching up. Cloud adoption and advanced analytics rates in SMEs remain below the EU average. The Polish Economic Institute indicates that in 2025 only **8.7% of companies used AI** - 2026 requires a massive educational effort. ### Cybersecurity in a geopolitical context Due to its geopolitical position, Poland remains on the front line of cyberwarfare. In 2026, experts predict intensification of attacks on critical infrastructure and AI-driven disinformation campaigns. **Protecting digital borders is becoming as important as protecting physical ones**. ## Practical recommendations: what to do in 2026 After analyzing hundreds of pages of reports and forecasts, I draw five key recommendations for technology leaders and decision makers: ### 1. Agentic readiness audit Don't force AI adoption. Instead: - Identify processes suitable for autonomization (repetitive, clearly defined, measurable) - Prepare data infrastructure (Data Governance, data quality, availability) - Define control points and escalation mechanisms to humans - Set KPIs for AI deployments - without measurable ROI, the project makes no sense ### 2. Infrastructure: hybrid, not monolithic - **Cloud**: flexibility for training and experiments - **On-premises**: stability for critical inference and data sovereignty compliance - **Edge**: ultra-low latency for real-time applications Consider equipping employees with AI PCs - in the long run cheaper than cloud subscriptions billed per query. ### 3. Compliance as competitive advantage Treat the EU AI Act not as an obstacle but as a **framework for building a secure business**. Transparency will attract clients tired of disinformation and AI unpredictability. - Start inventorying AI systems in your organization - Classify them by risk level (AI Act categories) - Implement technical documentation and decision logging mechanisms - Consider using the regulatory sandbox ### 4. Model specialization over universality If you operate in a regulated industry (finance, healthcare, law), **invest in DSLMs instead of fighting with general-purpose LLMs**. Hybrid architecture (foundation model + domain modules) is the sweet spot between flexibility and precision. ### 5. Reskilling as strategy, not tactics Especially relevant in Poland: - Invest in retaining and developing employees 50+ - their experience + new AI tools is key to stability - Create upskilling programs for teams - AI won't replace experts, but experts without AI will be replaced by experts with AI - Build a culture of experimentation and learning in your organization ## Summary: we're building foundations, not toys 2026 is the time when technology stops being "magic" and becomes "engineering." Fascination with AI's capabilities gives way to the hard work of deploying, securing, and scaling it. **Key takeaways**: 1. **Agentic AI** redefines automation - from supporting tools to autonomously acting systems 2. **Model specialization (DSLMs)** beats universality in regulated industries 3. **Hybrid infrastructure** (cloud + on-premises + edge) is the new standard 4. **EU AI Act** (August 2026) enforces transparency and accountability 5. **ROI and value verification** become critical - the era of experiments without strategy is over 6. **Physical AI** enters production - humanoid robots, autonomous factories 7. **Preventive cybersecurity** and AI Security platforms are a necessity 8. **The labor market** evolves - new roles, redefined competencies, reskilling Those who win will be the ones who, instead of waiting for the dust to settle, start building the foundations of the new, autonomous reality today. And if you're asking where to start - **start with an audit**. Check where in your organization AI can deliver measurable value, which processes are suitable for autonomization, what data you have available, and whether your infrastructure is ready. Because 2026 won't be about who has the best AI presentation. It will be about who has the best deployment.

Need support with AI implementation?

I'll help you assess your organization's readiness, identify processes for automation, and plan your first steps.

Book a free consultation
**Sources and references**: This article was created based on analysis of reports from Gartner, Forrester, IDC, Deloitte, McKinsey, Cisco, and EU AI Act documentation. All forecasts and data are current as of January 1, 2026. ## FAQ
### How does Agentic AI differ from generative models like ChatGPT or Claude? Agentic AI doesn't just answer questions - it autonomously plans and executes tasks in business systems. Instead of waiting for step-by-step instructions, agents independently make decisions and use tools, acting as "virtual workers" rather than just assistants.
### What obligations does the EU AI Act impose on companies from August 2026? From August 2, 2026, companies must implement risk management systems, human oversight, and complete technical documentation for high-risk AI systems. Training data quality assurance and incident reporting procedures are required, with penalties for non-compliance reaching 35 million euros or 7% of turnover.
### Why are Domain-Specific Language Models (DSLMs) better for business than general LLMs? DSLMs offer higher precision on specialist tasks (e.g., law, healthcare) at significantly lower operational costs thanks to fewer parameters. They also guarantee better regulatory compliance and data security, which is critical in regulated industries where general models often "hallucinate."
### Will AI in 2026 replace jobs or change their nature? AI in 2026 doesn't eliminate professions wholesale but redefines roles, automating routine cognitive tasks. New positions like AI Orchestrator and AI Risk Officer are emerging, and the key to maintaining employment is reskilling and the ability to collaborate with digital agents.
### What is the advantage of Reasoning Models over traditional language models? Reasoning Models (System 2) spend additional compute time on "deliberation," simulation, and hypothesis verification before providing an answer. This allows solving complex logical and mathematical problems with accuracy above 90%, at the cost of longer response times and higher token consumption.
--- # Vibe Coding - how to create UI with AI without design skills Source: https://pawel.lipowczan.pl/en/blog/vibe-coding-guide Published: 2025-12-24 # Vibe Coding - how to create UI with AI without design skills A few weeks ago I came across a piece by the PageAI team about vibe coding and I have to say - they did an excellent job. What's more, their approach perfectly matches how I personally work with AI when creating interfaces. If you don't have a UX/UI background (like me) but want to create modern, attractive interfaces - this guide is for you. ## What is Vibe Coding? Vibe coding is an approach to building interfaces where instead of precise mockups and pixel-perfect designs, you **describe the feeling and vibe** you want to achieve. AI (combined with good tooling) translates that into working code. It's the democratization of UI creation. AI handles reproducing ready-made templates, designs from a photo, or text descriptions remarkably well. Even without design skills you can create something that looks professional. ## 3 Pillars of Effective Vibe Coding From experience I know that vibe coding works when you have three things: ### 1. A Good Base (Starter Kit) Don't start from scratch. Use a proven boilerplate that has: - Configured tools (Vite, React, Tailwind) - A base design system (colors, typography, spacing) - Basic components **Recommended:** [Vibe Coding Starter Kit by PageAI](https://github.com/PageAI-Pro/vibe-coding-starter) It's a solid base with React + Vite + Tailwind + shadcn/ui. Everything ready for a quick start. ### 2. Good and Proven Prompts This is the heart of vibe coding. You need to know **how to talk to AI** to get sensible results. Below you'll find proven prompts you can copy and adapt. ### 3. Good Context for AI AI needs to understand: - What your design system is (colors, fonts, spacing) - What libraries you use (Tailwind, shadcn/ui, etc.) - What style you're designing (minimalist, bold, glassmorphic) Without context you get a generic, "AI-looking" design. With context - something unique. ## Step 1: Define the Design Brief Before you ask AI for code, you need to know **what you want**. Use this prompt: ```text I need help creating a design brief for [type of project]. Please help me define: 1. Design Style & Aesthetic - What visual style best suits this project? - What mood/feeling should the design evoke? - Any specific design movements or styles to reference? 2. Target Audience - Who will use this? - What are their preferences and expectations? - What devices will they primarily use? 3. Key Features & Priorities - What are the core features to highlight? - What's the primary user action? - What should stand out most? Based on my answers, create a comprehensive design brief that I can use with AI coding assistants. ``` **Example usage:** Me: "I need help creating a design brief for a personal portfolio website." AI will help you think through the style (e.g., minimalist with neon accents), target audience (tech recruiters, freelance clients), and key elements (hero section, projects, contact). The result? A clear brief you can pass along in subsequent prompts. ## Step 2: Create an AI Design Style Reference This is where the magic starts. Create a document that AI will use as a reference point for **the entire project**. ### Visual Style Define the visual style in a table: | Aspect | Choice | Rationale | |--------|-------|--------------| | **Overall Aesthetic** | Minimalist Brutalism | Highlights content, modern vibe | | **Color Palette** | Monochromatic + neon accent | Readable, expressive, tech-forward | | **Typography** | Inter (sans-serif) | Readable, professional | | **Spacing** | Generous whitespace | Eases focus, premium feel | | **Visual Weight** | Bold typography, subtle UI | Content > decoration | ### Layout Structures | Layout type | When to use | Characteristics | |-------------|--------------|-----------------| | **Hero Full-Screen** | Landing pages | Maximum impact, CTA above the fold | | **Sidebar Layout** | Dashboards, applications | Navigation always visible | | **Card Grid** | Portfolios, galleries | Scannable, responsive | | **Split Screen** | Comparisons, dual content | 50/50 attention split | ### Color Themes | Theme | Primary | Background | Text | Accent | |-------|---------|------------|------|--------| | **Dark Mode** | `#0f172a` | `#020617` | `#f1f5f9` | `#22d3ee` | | **Light Mode** | `#ffffff` | `#f8fafc` | `#0f172a` | `#0ea5e9` | | **Neon Dark** | `#000000` | `#0a0a0a` | `#ffffff` | `#00ff9d` | **Pro tip:** Save these tables as a separate file (e.g., `design-reference.md`) and include them in the AI context (Cursor, Claude, ChatGPT). ## Step 3: Implement with AI Now you have a brief and style reference. Time for code. Use this prompt: ```text I need you to build [specific component/page] using React and Tailwind CSS. DESIGN CONTEXT: [Paste your design brief and style reference] TECHNICAL REQUIREMENTS: - Use React functional components with hooks - Use Tailwind CSS for all styling (no custom CSS) - Make it fully responsive (mobile-first) - Follow accessibility best practices (WCAG 2.1 AA) - Use semantic HTML SPECIFIC REQUIREMENTS FOR THIS COMPONENT: - [Describe exactly what it should do] - [What sections/elements to include] - [Any special interactions] VISUAL DETAILS: - Follow the [specific style] from the design reference - Use [specific color theme] - Ensure [specific spacing/typography rules] Please provide: 1. Complete, production-ready component code 2. Brief explanation of design decisions 3. Any additional dependencies needed Start with the code, then explain. ``` ### Real-world usage example: ```text I need you to build a Hero section for a SaaS landing page using React and Tailwind CSS. DESIGN CONTEXT: - Style: Minimalist with gradient background - Color: Dark mode with cyan accent (#22d3ee) - Typography: Inter, bold headlines - Spacing: Generous (p-8, gap-8) TECHNICAL REQUIREMENTS: - React functional component - Tailwind CSS only - Mobile-first responsive - WCAG 2.1 AA compliant SPECIFIC REQUIREMENTS: - Headline + subheadline + CTA button - Gradient background (slate to cyan) - CTA button with hover effect - Centered content, max-width container VISUAL DETAILS: - Use Dark Mode theme from reference - Cyan accent for CTA - Bold typography (font-bold, text-5xl for headline) Start with the code. ``` AI will give you a ready component you can use right away. ## Design Tokens - Your Design System in JSON One of the best practices is keeping design tokens in JSON. This gives AI a **single source of truth** about your colors, spacing, and typography. Example `design-tokens.json`: ```json { "colors": { "primary": { "50": "#f0fdfa", "100": "#ccfbf1", "500": "#14b8a6", "900": "#134e4a" }, "neutral": { "50": "#f8fafc", "900": "#0f172a" }, "accent": { "cyan": "#22d3ee", "neon": "#00ff9d" } }, "typography": { "fontFamily": { "sans": ["Inter", "system-ui", "sans-serif"], "mono": ["Fira Code", "monospace"] }, "fontSize": { "xs": "0.75rem", "sm": "0.875rem", "base": "1rem", "lg": "1.125rem", "xl": "1.25rem", "2xl": "1.5rem", "3xl": "1.875rem", "4xl": "2.25rem", "5xl": "3rem" }, "fontWeight": { "normal": "400", "medium": "500", "semibold": "600", "bold": "700" } }, "spacing": { "xs": "0.5rem", "sm": "1rem", "md": "1.5rem", "lg": "2rem", "xl": "3rem", "2xl": "4rem" }, "borderRadius": { "none": "0", "sm": "0.25rem", "md": "0.5rem", "lg": "1rem", "full": "9999px" } } ``` **How to use it?** 1. Create a `design-tokens.json` file in your project 2. Import it in your Tailwind config (`tailwind.config.js`) 3. Include it in prompts: "Use colors and spacing from design-tokens.json" AI will stick to your design system. ## Tools I use ### Cursor (my #1 choice) Cursor is an IDE based on VS Code with built-in AI. It supports: - **Composer** - creates/modifies multiple files simultaneously - **Chat** - conversation with full project context - **Cmd+K** - inline code editing **Why Cursor?** Because AI sees the entire project. It knows what components you use, what design system you have, what folder structure. It generates consistent code. ### Shadcn/ui A component library you can install via CLI. The components are **yours** - you copy them into your project and modify them. AI knows shadcn/ui really well, so when you say "use shadcn Button component," you'll get exactly that. ### v0.dev (optional) A Vercel tool for generating UI from text prompts. Good for quick prototypes, but the code often needs polishing. ## Practical tips ### 1. Start with small components Don't try to generate an entire page right away. Start with: - Button - Card - Input - Navigation Once you have base components, AI can compose them into larger layouts more easily. ### 2. Iterate on live code Vibe coding isn't "generate and done." It's iteration: 1. Generate the first version 2. Preview in the browser 3. Refine the prompt ("Make the button larger, add hover effect") 4. Generate again 5. Repeat ### 3. Use screenshots as input AI (especially Claude with vision) can reproduce a design from a photo. Find inspiration on Dribbble/Behance, take a screenshot, paste it into AI: ```text Recreate this design using React and Tailwind CSS. Maintain the layout and color scheme, but adapt it to our design system. ``` It works surprisingly well. ### 4. Keep documentation alongside Include in context: - `design-reference.md` - Your style guide - `design-tokens.json` - Design system - `component-patterns.md` - How you use your components AI will stay consistent with your project. ## Pitfalls to avoid ### Generic "AI design" Without context AI generates boring, generic UI. Always: - Specify a concrete style (minimalist, brutalist, glassmorphic) - Point to inspirations (e.g., "like Stripe's homepage") - Define colors and typography ### Components that are too large Don't ask AI for "an entire landing page." Break it into: - Hero section - Features section - Pricing section - Footer Smaller pieces = better control. ### No design system If every component has different shades of blue and different border radii, your UI looks like Frankenstein. Create design tokens and stick to them. ## Key takeaways 1. **Vibe coding works** - even without design skills you can create professional UI 2. **The 3 pillars are key** - good base, good prompts, good context 3. **AI is a tool, not a magic wand** - you need a clear brief and iteration 4. **A design system is the foundation** - create it once, use it everywhere 5. **Get inspired, don't copy** - AI helps you create something unique, not generic ## Resources for further learning - [Vibe Coding Starter Kit](https://github.com/PageAI-Pro/vibe-coding-starter) - solid base to start - [Tailwind CSS Docs](https://tailwindcss.com) - official documentation - [Shadcn/ui](https://ui.shadcn.com) - component library - [Realtime Colors](https://realtimecolors.com) - color palette generator - [Font Pair](https://fontpair.co) - typography inspiration ## What to do next? If you want to get started with vibe coding: 1. **Clone the starter kit** from GitHub 2. **Create a design brief** for your project (use the prompt from this article) 3. **Define a style reference** - tables with visual style 4. **Generate your first component** (start with Button) 5. **Iterate and learn**

Need support implementing AI in product development?

I'll help you configure AI tools, automate workflows, and implement vibe coding in your project. From tool selection through environment setup to developer process optimization.

Book a free consultation
## FAQ
### What is vibe coding and who is it for? Vibe coding is an approach to UI creation where instead of pixel-perfect mockups you describe the feeling and style you want to achieve. AI translates that into working code. It's designed for people without a UX/UI background who want to create professional interfaces - developers, entrepreneurs, product creators.
### What are the 3 pillars of effective vibe coding? A good base (starter kit with Vite, React, Tailwind, and basic components), good prompts (specific instructions for AI with design context), and good context (design system, libraries, project style). Without these three elements AI generates generic, inconsistent designs.
### Why keep design tokens in a JSON file? Design tokens are a single source of truth about your project's colors, spacing, and typography. When AI includes this file in context, it maintains a consistent design system. Without tokens, every component has different shades and border radii - the interface looks like Frankenstein assembled from random parts.
### What tools are best for vibe coding with AI? Cursor (IDE with built-in AI that sees the whole project), shadcn/ui (components copied into your project, well known by AI), and optionally v0.dev (Vercel for quick prototypes). Cursor is the number one choice - AI generates consistent code because it knows the folder structure and existing components.
### How to avoid generic "AI design" when vibe coding? Always specify a concrete style (minimalist, brutalist, glassmorphic), point to inspirations ("like Stripe's homepage"), and define colors and typography in a design reference. Break tasks into small components instead of asking for an entire page at once. Iterate - vibe coding is a process of generating, previewing, and refining, not a one-shot generation of finished UI.
--- # Data as Business Fuel - From Excel to AI Source: https://pawel.lipowczan.pl/en/blog/data-as-business-fuel Published: 2025-12-21 Data is the new oil. This statement has been circulating in the tech industry since 2006, when mathematician Clive Humby first popularized the comparison. And just like crude oil, data has no value on its own -- only after proper "refining" does it become fuel that powers business. From my 15 years of experience as a developer and 4 years working with no-code, I know one thing: **most companies are sitting on a gold mine but can't extract it**. They have tons of data but lack the tools and processes to turn it into valuable insights. And now, in 2025, we're on the brink of another revolution. AI agents like Claude Code, GitHub Copilot, and Cursor are changing the rules of the game -- what used to require months-long IT projects can now be built in days. Combining no-code agility with the power of traditional code backed by AI opens up entirely new possibilities. Today I'll show you how to make that journey -- from chaos in spreadsheets to an organized data system that actually works for the business. ## The Problem: Drowning in Data, Dying of Thirst for Information ### 180 Zettabytes of Chaos Imagine **180 zettabytes** of data. That's the amount of information that will be generated in 2025. To put this in perspective: if we printed that data on A4 sheets and stacked them, we could fly to the Moon and back **60 times**. Or we'd need **62 billion 16GB USB drives**. The amount of data doubles every 3 years. And it will likely double even faster due to the growth of AI and IoT. **The paradox is this**: we have plenty of data but can't extract information from it. It's like sitting in front of a pantry full of ingredients but without a chef or a cookbook -- we can't make a pizza. ### Four Key Challenges From my experience working with dozens of companies, I see the same problems repeating like a mantra: **1. Data Silos** Customer data in one system, orders in another, invoices in a third. No integration. Want to see the full customer picture? You need to cobble together 3-4 different reports and manually merge them in Excel. **2. Poor Quality and Inconsistency** - Duplicates (the same customer entered 5 times) - Typos (Smith, Smyth, Smiith) - Different formats (2025-12-21 vs 12/21/2025 vs 21.12.25) - Wrong data (someone entered a birth date instead of an order date) Any analysis based on such data leads you astray. **3. Difficult Access** Data sits in SQL databases that regular employees can't access. Or they can access them but can't write a query. Every simple report requires a ticket to IT and a week of waiting. **4. Lack of Data Culture** Even when we give people tools, they don't know: - What questions to ask - How to interpret results - How to draw conclusions - How to implement those conclusions Research from 2020 shows that despite having access to data, companies: - Lack a shared picture of the situation - Postpone decisions (because they're "not sure") - Waste precious time searching for information - Squander the economic and financial value of their data ## The Path from Raw Material to Value Before data can contribute anything valuable, it must go through several key stages: 1. **Acquisition** - collecting data from various sources 2. **Integration** - combining different sources into one coherent picture (this is the most important step!) 3. **Cleaning and transformation** - without this, analyses are based on false premises 4. **Storage** - in an analytical repository, accessible to the right people 5. **Analysis and visualization** - dashboards, reports, charts 6. **Implementing insights** - actually using information in decision-making **This is where business value finally appears.** Most companies get stuck at steps 2-3. They have data but can't merge and clean it well enough to serve as a basis for decisions. ## Solution Part 1: No-Code and Data Democratization Over the past 4 years, I've watched no-code transform how companies work with data. Platforms like **Airtable, Make, n8n, and Zapier** radically lower the barrier to entry. People from business, marketing, and finance departments can independently: - Create data structures - Integrate different sources - Build simple reports and dashboards - Automate information flow They don't have to wait weeks for the IT department. They don't need to know SQL, Python, or JavaScript. ### Benefits of Data Democratization **Speed**: A prototype system in Airtable can be built in hours, not months. **Cost**: No need to engage expensive developers for simple reports. **Business proximity**: The people who know the data best (because they use it) can build solutions themselves. **Engagement**: Employees feel empowered -- they can solve their own problems. ### But Democratization Requires Balance Data democratization isn't just about technology. It's two conflicting values that must coexist: **Data literacy** - user skills: - Effective use of data - Asking the right questions - Drawing sensible conclusions - Understanding limitations and errors **Data governance** - data management: - Not everyone should have access to everything - There must be procedures and security frameworks - Quality and consistency must be maintained - Someone needs to make sure data doesn't get "broken" You need to find a **golden mean** -- governance can't be a brake, but it must be agile enough to support users rather than block them. ## Solution Part 2: AI as a Game-Changer And this brings us to where we are now -- late 2025. **AI is changing absolutely everything when it comes to working with data.** ### What AI Already Delivers Today **1. Talking to Data in Natural Language** You don't need to know SQL. You type: "Show me the top 10 customers by order value this quarter" -- and you get an answer. This radically lowers the data literacy barrier. **2. Data Cleaning and Normalization** AI can find duplicates even with typos. It can standardize formats. It can fill in gaps based on context. **3. Pattern Discovery** In large datasets, AI will find correlations that humans would miss. For example: "customers who buy product X on Friday are more likely to return for product Y on Monday." **4. Generating Analyses and Forecasts** Based on historical data, AI can predict future trends, demand, and customer churn risk. **5. Report Automation** Instead of manually creating reports every week, AI can automatically generate summaries, charts, and insights. ### But Also Risks We need to be aware of the limitations: - **Hallucinations** - AI can invent facts that sound convincing - **Black box** - we don't always know where AI got its information - **Non-determinism** - it may give a different answer each time - **Garbage in, garbage out** - junk data will produce junk insights - **Privacy** - data can leak to models trained by third-party companies - **Bias** - even well-organized data can lead to wrong conclusions due to model biases That's why AI is a tool, not a replacement for thinking. ## The Breakthrough: AI Agents for Developers And now the most important part: **a new generation of AI tools for developers is changing the entire game**. For years we had a choice: - **No-code**: fast, cheap, but limited capabilities - **Traditional code**: unlimited capabilities, but slow and expensive ### AI Agents Bridge These Two Worlds Tools like: - **Claude Code** (which I'm using to write this article) - **GitHub Copilot** - **Cursor** - **Windsurf** ...give developers **no-code agility + the power of traditional code**. What was previously only possible through no-code (rapid prototyping, iterative solution building), developers with AI assistance can now do just as quickly -- but with the full flexibility of code. ### A Real-World Example I recently built a system for analyzing data from multiple sources (Airtable, external APIs, CSV files). Previously: - In no-code: fast, but I couldn't handle complex transformations - In traditional code: 2-3 weeks of work With Claude Code: **3 days**. AI helped me: - Quickly write API integration scripts - Generate data cleaning code - Build an ETL pipeline - Write tests - Optimize queries I didn't need to memorize every library's syntax. I didn't need to search for solutions on Stack Overflow. I simply described what I wanted to achieve -- and AI generated code I could review, understand, and customize. ### What This Means for Data Management **1. Faster Solution Deployment** Projects that used to take months now take weeks. This means faster ROI and the ability to experiment. **2. Lower Barriers to Entry - But With a Caveat** AI lowers the entry threshold, but **it doesn't eliminate the need for experience**. A junior developer with AI can do more, but an AI agent on its own is like a junior -- it can make mistakes, miss business context, and generate code with security gaps. **Experience is not a requirement, but it's strongly recommended.** Someone experienced needs to oversee what AI generates -- verifying approaches, catching errors, assessing quality. AI is a powerful tool, but it requires supervision. **3. Better Code Quality** AI helps with writing tests, documentation, and optimization. Code is cleaner and easier to maintain. **4. Code vs No-Code - Depends on Your Team** The choice depends on the resources in your organization: **Have a senior developer?** -> Go with code + AI. You'll get flexibility, no limitations, and full control. **No technical team?** -> Stick with no-code (Airtable, Make, n8n). It's a far more accessible solution for people without programming experience. **A no-code + code hybrid** only makes sense if you have someone technically strong on board. Then it's the best option: you keep the flexibility of code, the agility of prototyping, and practically no limitations that sometimes appear in no-code tools. **5. Talking to Data in Natural Language - But With the Power of SQL** You don't have to choose between simplicity and flexibility. AI can translate your question into a complex SQL query or Python script. ## Case Study: From Chaos to Clarity Theory is one thing, practice is another. Let me show you how we went through this transformation in our own company. ### Context 22Ventures was a holding company made up of several firms: Automation House, Tigers, Huciao, Sowicki Legal. Automation House was created specifically to organize processes and data internally -- and then broadly bring those solutions to market. We started where most companies do: **chaos in spreadsheets**. ### Problem 1: Data Inconsistency Every team had their own spreadsheets. The same data was entered differently: - Customer duplicates (the same customer 5 times in the database) - Typos (Smith, Smyth, Smiith) - Different date, amount, and name formats - Transcription errors **Solution**: Normalization in Airtable - Creating dictionaries (lists of unique values) - Removing duplicates - Standardizing formats - Field-level validation (only specific values, formats) ### Problem 2: Multiple Data Sources Each department had its own spreadsheet: - HR - employee list - Finance - invoices and payments - Projects - tasks and timesheets - Sales - leads and customers Data was duplicated across departments. No synchronization procedures. No organizational culture around data. **Solution**: Migration to Airtable Smooth migration from spreadsheets to Airtable. Consolidation into one workspace with proper bases and relations. **Example transformation**: Team leader **In a spreadsheet**: Every project had a full row of team leader data: first name, last name, email, phone, rate. Changing team leader data = manual changes in 20 places. **In Airtable**: An employees base + a projects base connected by a relation. The team leader is a linked record. Change the data in one place -- it works everywhere automatically. ### Problem 3: No Dictionaries Data was entered freehand. Everyone typed it their own way. **Example**: Employee benefits In a spreadsheet: Every employee had a "Benefits" column with a full description and price. Changing the gym membership price = manual change for 50 employees. **In Airtable**: A benefits base (dictionary) with prices. Employees have linked records to benefits. Change the price in one place -- it updates automatically for everyone. ### Problem 4: Redundancy and Duplication Spreadsheets naturally lead to data duplication. Every change requires manual propagation. **Solution**: Single source of truth In Airtable, every record exists in one place. Relations connect data without duplication. Change in one place = change everywhere. ### Transformation Results **Before**: - 15 different spreadsheets - 3-4 hours per week preparing reports - Data errors in every analysis - Team frustration - Decisions based on "gut feeling" **After**: - One system in Airtable - Reports generated automatically - Clean and consistent data - Happy team (they have tools that work) - Data-driven decisions **And most importantly**: This system evolves. We started at level 1-2, and now we're at 3-4 out of my 5 data maturity levels. ## 5 Levels of Data Maturity Based on working with dozens of companies, I've developed a model for assessing where an organization stands in data management: ### Level 1: Ad Hoc - Spreadsheets scattered across the company - First attempts at analysis (basic SUM, AVERAGE formulas) - Everyone does it their own way - No standards ### Level 2: Consolidation - Combining data sources into one place - E.g., migrating from spreadsheets to Airtable - Everything in one system - Basic relationships between data ### Level 3: Standardization - Data collection procedures - Processing standards - Data governance (who has access to what) - Dictionaries and validation - Data cleaning processes ### Level 4: Optimization - Financial utilization of data - Creating data-based products - Monetization (selling data, insights) - Advanced analyses and forecasts ### Level 5: Innovation - Fully data-driven culture - Every decision based on facts - Standardized processes across the company - Regular monetization - Continuous improvement of data quality and flow - Experimentation and hypothesis testing **Most companies are at level 1-2. Very few reach 4-5.** But with no-code and AI, this process can be accelerated many times over. ## What You Can Do Today ### Step 1: Assess Your Maturity Ask yourself: - How many systems/spreadsheets do we store data in? - How long does it take to prepare a basic report? - How often do we find errors in our data? - Do we have procedures for data entry? - Do people trust our data? This will help you determine which level you're at. ### Step 2: Start with One Area Don't try to fix everything at once. Pick one problem: - Maybe it's a customer list with duplicates - Maybe it's project chaos - Maybe it's the sales reporting process Start with a small pilot project. ### Step 3: Choose the Right Tool **For simple cases**: Airtable, Notion, Google Sheets with Apps Script **For automation**: Make, n8n, Zapier **For advanced analytics**: Python + Pandas (with AI assistance), Power BI, Tableau **For no-code + code integration**: Airtable + custom scripts built with Claude Code or Cursor ### Step 4: Build a Prototype with AI You don't need to be an expert. Use Claude, ChatGPT, or another LLM: - Describe your problem - Ask for a data structure proposal - Ask for integration/cleaning code - Iterate and refine With AI, you can build a working prototype in hours, not weeks. ### Step 5: Test, Learn, Iterate Don't seek perfection right away. Deploy something simple, see how it works, gather feedback, improve. **Agility is key.** ## Key Takeaways ### 1. Data Without Refining Is Just Junk It's not enough to have lots of data. You need to merge, clean, and standardize it -- only then does it have value. ### 2. No-Code Democratizes Access Airtable, Make, n8n -- these tools give business people the power to build solutions without waiting for IT. ### 3. AI Changes the Rules of the Game Talking to data in natural language, automatic cleaning, pattern discovery -- AI lowers the barrier to advanced analytics. ### 4. AI Agents Bridge No-Code and Traditional Code Claude Code, GitHub Copilot, Cursor -- now developers can build as agilely as with no-code, but with the full power of code. It's the best of both worlds. ### 5. Democratization Requires Balance Data literacy (skills) + data governance (management) must go hand in hand. Easy access without control is chaos. Control without access is a brake. ### 6. Start with a Small Step Don't wait for a grand transformation project. Pick one problem, build a prototype, test, learn. Small wins build momentum. ### 7. Quality > Quantity It's better to have 10 well-organized, clean data sources than 100 chaotic spreadsheets. ## Next Steps The world of data is changing at lightning speed. What was science fiction a year ago (AI generating code, talking to data in natural language) is now an everyday reality. **Companies that master the art of turning data into business value will be leaders in their industries.** And technology is no longer the barrier. No-code lowered the entry threshold. AI provided analytical power. AI agents gave developers agility. **Now the only barrier is the decision: start or wait.** From my experience, I know that companies that started a year ago have a competitive advantage today. Those that start today will have one a year from now. And those that keep waiting... will drown in a sea of data, dying of thirst for information.

Need help organizing your company's data?

I can help you go from spreadsheet chaos to a coherent data management system. From auditing and consolidating sources through migration to Airtable/databases to automating reporting.

Book a free consultation
## FAQ
### What does "data is the new oil" actually mean in business practice? Data, like oil, has no value on its own -- it must go through "refining" (cleaning, integration, analysis) to become fuel for business decisions. Most companies are sitting on a gold mine but can't extract it -- they have tons of data but lack the processes and tools to transform it into useful information.
### What are the main obstacles preventing companies from leveraging their data? Four key challenges: data silos (information scattered across different systems with no integration), poor quality (duplicates, typos, inconsistent formats), difficult access (needing IT for every report), and lack of data culture (people don't know what questions to ask or how to interpret results).
### When should you choose no-code vs AI-assisted code for data management? Have a senior developer on your team -- choose code with AI (Claude Code, Cursor) for full flexibility. No technical team -- stick with no-code (Airtable, Make, n8n). A hybrid approach only makes sense with strong technical support -- then you combine prototyping agility with practically no limitations.
### How do you assess your company's data maturity level? Five levels: (1) Ad hoc - scattered spreadsheets with no standards, (2) Consolidation - data in one system, (3) Standardization - procedures, governance, validation, (4) Optimization - monetization and advanced analytics, (5) Innovation - a fully data-driven culture. Most companies are stuck at level 1-2.
### Where should you start with a data transformation in your company? Pick one specific problem (e.g., customer duplicates, project chaos), build a small prototype in Airtable or with AI, test it with your team, gather feedback, and iterate. Don't seek perfection right away -- small wins build momentum. Trying to transform the entire company at once is a recipe for failure.
--- # Airtable vs Excel - When Is It Worth Switching from Spreadsheets to a Database? Source: https://pawel.lipowczan.pl/en/blog/airtable-vs-excel-migration Published: 2025-12-19 # Airtable vs Excel - When Is It Worth Switching from Spreadsheets to a Database? For years, Excel was synonymous with data organization in businesses. Everyone knows it, everyone uses it. But from my own experience, I know that at a certain point, classic spreadsheets stop being enough. When the team grows, data gets complex, and processes demand better collaboration - that's when frustration sets in. That's exactly why more and more companies are deciding to **migrate excel to airtable**. And this isn't about any tech snobbery - it's simply a practical solution to real problems that Microsoft Excel has baked into its DNA. ## The Problem: Limitations of Classic Spreadsheets Before we get to the solutions, let's take an honest look at what teams using Excel struggle with: **1. Versioning Hell** You know that feeling? `Projects_v2.xlsx`, `Projects_v2_final.xlsx`, `Projects_v2_final_TRULY_THE_LAST_ONE.xlsx`. Emails flying back and forth, nobody knows which version is current, and everyone's working on their own file. The result? Chaos, duplicated work, and errors. **2. Relationships Between Data Are a Nightmare** Let's say you're running a content calendar. You have articles, authors, marketing campaigns, publication statuses. In Excel? A mountain of VLOOKUPs, linked sheets that break at the slightest structural change. Manual data copying. Risk of errors at every step. **3. Real-Time Collaboration? Forget It** Yes, I know - there's Excel Online and Google Sheets. But honestly, that's still not true collaboration. No permission controls, no change history, no flexible views for different team roles. **4. Visualization and Reporting Require Gymnastics** Want to see projects on a Kanban board? A deadline calendar? A gallery with thumbnails? In Excel, that's either a macro, a separate dashboard, or... you just can't do it conveniently. For years, I built complex spreadsheets for clients and every time I reached a point where I thought: "There has to be a better way." And that's when I discovered the **airtable database builder**. ## The Solution: Airtable as a Relational Database with a Spreadsheet Interface ### What Is Airtable, Exactly? Put simply: **Airtable is a relational database that looks and feels like a spreadsheet**. That's the key difference when comparing **airtable vs excel**. Under the hood, it's a real database with relationships, field types, and data integrity. But on the surface? An intuitive interface that doesn't require knowledge of SQL or programming. ### Relationships Between Tables - Game Changer In Airtable, you can create: - A "Blog Articles" table - An "Authors" table - A "Marketing Campaigns" table And then **connect them to each other**. One click, a drag, and you can see all articles by a given author. All materials tied to a campaign. No VLOOKUPs, no formulas, no risk of the structure falling apart. This isn't theory - I use this daily at automation.house. Our knowledge base of clients, projects, and processes lives in Airtable and saves us dozens of hours every month. ### Different Views for Different Roles This is one of my favorite features. You can view the same data as: - **Grid** - classic spreadsheet view - **Calendar** - perfect for deadlines and planning - **Kanban** - for project management in the Trello style - **Gallery** - great for portfolios, products, graphics - **Form** - for collecting data from external people Marketing looks at campaigns through Kanban. The content writer through the publication calendar. The manager through a filtered table. **Everyone sees what they need.** ## Why Companies Migrate from Excel to Airtable From conversations with clients and my own experience, five main reasons stand out: ### 1. True Team Collaboration Everyone works on the same base (that's the Airtable term) in real time. Changes are instant. You can comment on records, tag people, set reminders. This is no longer a tool for individual work - it's a team platform. ### 2. Automating Repetitive Tasks Airtable has a built-in automation system. Real-life examples: - When a project status changes to "Pending Approval" -> send a notification to the manager - When a deadline is 3 days away -> send an email to the responsible person - When a new record is added via a form -> create tasks in linked tables In Excel? You need VBA, macros, or external tools. In Airtable? You click and configure. ### 3. Integrations with the Rest of Your Stack Airtable connects beautifully with: - Slack (notifications) - Gmail (sending emails) - Zapier/Make (advanced workflows) - Google Calendar (event sync) - And hundreds of other tools This makes Airtable a central data hub for the entire company. ### 4. Permission Control and Security You can precisely define who has access to which table, who can edit, and who can only read. You can hide fields from certain roles. Change history shows who modified what and when. ### 5. Painless Scalability Start with a simple base of 3 tables. Then add more. Connect them with relationships. Add automations. Create public forms. Build interfaces for clients. Airtable grows with your needs, not against them - as is often the case with overgrown Excel spreadsheets. ## How to Move Data from Excel to Airtable Alright, you're convinced. But how do you actually **migrate excel to airtable**? ### Step 1: Prepare Your Data in Excel - Make sure each table has headers in the first row - Remove empty rows and columns - Separate data logically (if you have multiple "entities" in one sheet, consider splitting them) ### Step 2: Import into Airtable Airtable allows importing `.xlsx` and `.csv` files. It's simple: 1. Create a new base in Airtable 2. Click "Add or import" -> "CSV or TSV" 3. Upload the file 4. Airtable automatically recognizes column types (text, numbers, dates) ### Step 3: Refine the Structure This is where the magic happens. After import: - Change field types where Airtable got it wrong (e.g., change text to email, URL, phone) - Split data into separate tables (e.g., separate clients from projects) - Create **relationships between tables** using the "Link to another record" field - Remove duplicates and organize data ### Step 4: Create Views - Build Calendar views for dates - Kanban for statuses - Gallery for projects with images - Filtered views for specific teams ### Step 5: Add Automations Start with simple ones: - Slack notification when a new record is added - Email when a deadline is approaching - Automatic status change ## Airtable vs Excel - Who Is Airtable For? I'm not saying Excel is bad. It has its place. But Airtable is better for: **Teams collaborating on data** Excel is for individual work, Airtable is for teamwork. **Data with relationships and dependencies** Clients -> Projects -> Invoices -> Payments. In Airtable, that's natural. In Excel - painful. **Processes requiring different views** Calendar, Kanban, Grid, Gallery - all from the same data. **Automation and workflows** Airtable has this built in, Excel requires VBA or external tools. **Integration with other tools** API, Zapier, Make - Airtable is built for connecting with the ecosystem. On the other hand, **Excel still wins** for: - Advanced financial and statistical calculations - Individual data analysis - One-off reports - Offline work without internet access ## What Can You Do Today? If you're considering migration, start with small steps: 1. **Pick one process/spreadsheet** to test in Airtable 2. **Use the free plan** of Airtable (enough to get started) 3. **Import your data** and play with the structure for a week 4. **Check if it solves your Excel problems** 5. **Only then consider a full migration** From my experience, the best first candidates are: - Content calendars - Client databases/CRM - Project management - Team task lists - Inventory/product catalogs ## Key Takeaways The **airtable vs excel** comparison doesn't have a clear-cut winner - it depends on context. But if: - You work in a team - Data has complex relationships - You need different views of the same data - You want to automate workflows - You integrate data with other tools ...then the **airtable database builder** will likely save you dozens of hours per month and a lot of frustration. I went through this transition myself a few years ago and can't imagine going back to managing projects and client data in Excel. It's like switching from a Nokia 3310 to a smartphone - technically both are phones, but the capabilities are incomparable.

Need help migrating from Excel to Airtable?

I'll help you safely transfer your data, design the database structure, configure automations, and train your team. From needs analysis through migration to deployment and support.

Book a free consultation
## FAQ
### What is the main difference between Airtable and Excel? Airtable is a relational database with a spreadsheet interface, while Excel is a spreadsheet. In Airtable, you can create relationships between tables (e.g., Clients -> Projects -> Invoices) without VLOOKUPs, and data automatically stays in sync. Excel is great for individual work and calculations, Airtable for team collaboration.
### How do you migrate data from Excel to Airtable step by step? Prepare your data in Excel (headers in the first row, remove empty rows), import the .xlsx file into a new Airtable base, refine field types, and create relationships between tables. Finally, add views (Calendar, Kanban, Gallery) and configure automations. The entire process takes anywhere from a few hours to a few days, depending on data complexity.
### When is it better to stick with Excel instead of migrating to Airtable? Excel wins for advanced financial and statistical calculations, individual data analysis, one-off reports, and offline work without internet access. If you work solo on data without relationships and don't need automation - Excel is sufficient.
### What are table relationships in Airtable and why are they important? Relationships are connections between tables that let you link, say, articles to authors with a single click. Instead of VLOOKUPs that break with structural changes, Airtable automatically syncs related data. Clicking an author shows all their articles without manual filtering or data copying.
### What process should you start with when migrating to Airtable? The best starting points are: content calendars, client databases/CRM, project management, team task lists, and product catalogs. Pick one simple process, test it for a week on Airtable's free plan, check if it solves your Excel problems - and only then plan a full migration.
--- # Hackathon Hacknation - Experience Analysis and a Lesson in Digitalization Source: https://pawel.lipowczan.pl/en/blog/hackathon-hacknation-experience-analysis Published: 2025-12-12 # Hackathon Hacknation - Experience Analysis and a Lesson in Digitalization Hackathon Hacknation, organized by GovTech Poland, was a massive event. Over 1,500 participants, 480,000 PLN in the prize pool, and one goal: build a working solution for public administration challenges in 24 hours. For our team -- which I'd describe as junior-level programmers -- it was more than a competition. It was a testing ground. I'm a "retired" programmer. I stopped actively coding about 4 years ago, which in the tech world is practically light years -- tools and frameworks have changed so much that you essentially need to start from scratch. While software engineering experience is very helpful, without knowledge of current languages and environments, you're sometimes working blind. The other team members had never had much to do with traditional programming. In their day-to-day work, they use no-code technologies supported by language models. ![Hacknation Team](/images/hacknation-team.webp) Together, we confronted our expectations of "intelligent" AI agents with harsh reality. Here's the story of how technology met bureaucracy, why lack of validation can kill the best project, and what we learned about collaborating with AI under time pressure. ## 1. Mrs. Zosia and Thousands of Spreadsheets - Problem Analysis Our challenge involved the budgeting process in public administration. Sounds boring? Maybe, but the scale of the problem is enormous. The core problem turned out to be a process based on manually exchanging hundreds of thousands of Excel files. Errors, information chaos, lack of transparency -- that's daily life for government officials. We created a metaphor for this process, which we called **"Mrs. Zosia and Thousands of Spreadsheets"**: 1. **Start (Bottom):** "Mrs. Zosia" at the municipal office "reads tea leaves," manually entering budget data into Excel (e.g., a request for a new computer). 2. **Escalation (Top):** The file travels up the hierarchy: Municipal Office -> Regional Office -> Ministry of Finance. 3. **Consolidation:** A special unit at the ministry merges data from all files (often manually!). 4. **Decision and Return (Bottom):** Budget limits travel back the same way, often with arbitrary cuts. In the end, "Mrs. Zosia" learns she won't get the new computer, but nobody can explain why. ## 2. The Solution: Digital Budget We went with a simple but radical solution: **Digital Budget**. Instead of sending files around, let's move the entire process to the cloud. Our concept was built on a centralized web application with several key features: * **Single source of truth:** All budget items are entered in one system, visible (with appropriate permissions) to every level. * **Transparency and communication:** The ability to comment on and discuss each budget item directly in the system, instead of in emails. * **Approval workflow:** A simplified process for approving and consolidating the budget. Interestingly, we also employed AI to choose the challenge itself. We analyzed the available challenges against our team's competencies (mostly "non-programmers") to maximize our chances. The choice fell on budgeting, where understanding the business process seemed more important than complex algorithms. ## 3. The Atmosphere, Sweat, and Sleep Deprivation ![Hackathon Ending](/images/hacknation-end.webp) The atmosphere at Hacknation was incredible. A huge hall, open space, a stage, constant talks -- the energy of a thousand people "fired up about technology" was contagious. It was this vibe that kept us going even as fatigue grew with every hour. We worked non-stop for 24 hours. We slept 2-3 hours in the hallway or a dedicated sleeping room, where someone's phone alarm kept going off. Initially, each of us dove into work "freestyle," creating our own pieces of code. But we quickly realized that was a dead end. The turning point came when we decided to consolidate our efforts around Justyna's prototype, which was the most advanced. It became the foundation of our final solution. ## 4. AI as an "Equalizer" - Technology in Practice Our main thesis to test was: **AI is an equalizer**. A tool that allows a team with less coding experience (non-programmers) to compete with professional dev teams. Our technology stack: * **Frontend:** React, TypeScript * **Backend:** Supabase * **Presentation:** Video generated in HiGen We used heavy AI artillery: * **Pawel and Kuba:** Antigravity (Gemini Pro / Claude 4.5 models). We burned through the entire weekly token limit in a dozen hours. * **Justyna:** Bolt (Claude Code model). Record consumption of **18 million tokens**. ### A Contrarian Take on AI Did AI write the application for us? Not exactly. Despite the enthusiasm, we felt slightly disappointed. * **Code often didn't work:** AI generated solutions that looked correct but fell apart at runtime. * **Hallucinations:** Suggested libraries didn't exist, and pieces of logic were completely off. * **Hand-holding required:** Achieving a correct result demanded precise prompting and constant course correction. * **Blockers:** Justyna hit a bug in _current user_ filters that the model couldn't diagnose. We had to go back to basics -- reading code and debugging manually. AI is a powerful force multiplier, but not a magic wand. Without technical skills and critical thinking, we would have gotten stuck halfway. ## 5. The Biggest Weakness - Lack of Validation Our final score was **2.15 / 5 points**. We didn't make the finals. Why? The technology worked. The presentation was great. One crucial element was missing: **VALIDATION**. The team didn't have access to a practitioner -- a government official who works with budgets day-to-day. The mentor assigned to the challenge wasn't a domain expert. As a result, we built a system that seemed logical to us but could have been completely disconnected from administrative reality ("off-target"). This is the most important lesson: **Validation > Technology**. Even the best code can't save a solution that doesn't address real user needs. ## 6. Plan for the Next Hackathon Learning from experience, we prepared an improved process for the future: ### 1. Challenge Selection * Define team roles (strengths/weaknesses). * Scrape challenges and analyze them with an AI agent for team fit. * _Takeaway:_ We had this stage down pat. ### 2. Business Analysis (This Is Where We Failed!) * Prepare a current-state process map (AS-IS). * Design the target process (TO-BE). * Write User Stories. * **PRD (Product Requirements Document):** What are we building and why? * **SRS (Software Requirements Specification):** How are we building it? * Prepare "skills" for AI agents. ### 3. Development * Start with a prepared boilerplate (don't waste time on setup!). * Iterative development of functional requirements. * AI-generated automated tests. * Continuous Code Review. ### 4. Documentation and Verification * Review for security and performance. * Prepare documentation (optional). ## Summary and Takeaways Hacknation was an invaluable lesson for us. It confirmed that AI lets you do things that were impossible just a year ago -- a small team built a working web application in 24 hours. At the same time, it exposed a brutal truth: in the world of product development, **technology is secondary to understanding the problem**. Other teams that came with ready-made components and better business analysis homework won. The "freestyle" approach is romantic, but in a match against preparation -- it loses. AI is the future of programming, but it's still the human who must be the pilot who knows where they're flying.

Want to implement AI in your organization?

I can help you find real AI applications for your business, avoid common pitfalls, and deploy solutions that deliver measurable results. From concept through prototype to production.

Book a free consultation
## FAQ
### Can AI allow non-programmers to compete with professional teams at hackathons? AI is an equalizer -- a team without coding experience can build a working web application in 24 hours. However, AI isn't a magic wand: code often doesn't work, models hallucinate non-existent libraries, and achieving results requires constant course correction. Without technical skills and critical thinking, you'll get stuck halfway.
### What is the biggest mistake teams make at tech hackathons? Failing to validate the solution with the end user. You can build a technically working system that's completely disconnected from reality -- "off-target." Teams win not with the best code, but with better business analysis. Validation > Technology: even the best code can't save a solution that doesn't address real needs.
### How much does AI actually help when building a project in 24 hours? AI is a powerful force multiplier -- you can consume 18 million tokens and generate tons of code. But it requires hand-holding: proposed solutions look correct but fall apart at runtime. Debugging still requires manually reading code. AI accelerates development but doesn't eliminate the need for technical fundamentals.
### What preparation process increases your chances of winning a hackathon? Four stages: (1) selecting a challenge matched to team competencies, (2) business analysis with AS-IS/TO-BE process maps and user stories, (3) development with a prepared boilerplate and tests, (4) documentation and verification. The critical mistake is skipping stage 2 -- business analysis determines success more than code quality.
### Why is understanding the problem more important than technology at hackathons? Technology is secondary to understanding the user's problem. Teams with ready-made components and better business analysis win over teams with better code that didn't validate their solution. The "freestyle" approach is romantic, but in a match against preparation, it loses.
--- # Coding in 2025: Did AI Build My Portfolio? A Case Study of pawel.lipowczan.pl Source: https://pawel.lipowczan.pl/en/blog/coding-in-2025-ai-portfolio Published: 2025-12-02 I keep hearing that programming is dead. That all you need to do is "vibe code" an app in one of the new no-code tools and AI will handle the rest. I decided to put that to the test on a live project. I built [pawel.lipowczan.pl](https://pawel.lipowczan.pl) -- a project that was supposed to be a business card but turned into a testing ground for collaboration between an Experienced Engineer and an AI Agent. The verdict? If you think you can build a professional, secure, and scalable service without technical knowledge, just by "chatting" with a chatbot -- you're wrong. But if you have engineering fundamentals and treat AI as a junior developer on steroids, the results (and the costs) might surprise you. Here's a behind-the-scenes look at building my portfolio with a React + Vite + Tailwind stack. ![hero](/images/hero.webp) ## 1. The Foundation: Modern Stack and SEO in the SPA World My goal was simple: go beyond a static CV. I wanted aesthetics, performance, and a place to share knowledge. The choice fell on **React 18 + Vite**. Why? Because Vite delivers blazing-fast builds. However, React typically means a Single Page Application (SPA), which can be problematic for SEO. This is where engineering comes in. I implemented a **prerendering** mechanism. Even though React and Tailwind CSS run under the hood, we serve static HTML files to search engine crawlers. The result? The site is lightning fast, and Google sees it like a classic document. I host everything on **Vercel**, which turned out to be a bullseye. Built-in analytics and Core Web Vitals analysis let me hit "green scores" almost immediately after deploy. ![speed_insights](/images/speed_insights.webp) ![web_analytics](/images/web_analytics.webp) ## 2. Design for a Non-Designer: No More Trial-and-Error "Bells and Whistles" Let's be honest: I'm not a designer. I've always struggled with choosing color palettes, laying out elements, and animations. In a traditional workflow, I would have spent hours pushing pixels in CSS. That problem disappeared here. I could show AI examples of sites I liked, and the agent adapted that style to my project. Instead of experimenting with CSS code, I pointed to spots that needed fixing in the IDE, and AI corrected the layout in seconds. **Lovable vs. IDE** I experimented with various "vibe coding" tools, including Lovable, Vercel, and Firebase Studio. Lovable produced great visual results, but I ultimately chose to generate code directly in the IDE (Cursor). Why? Because I wanted full control and a modern, clean end result that's "my" code, not a closed black box. ## 3. Graphics: Consistency Over Perfection In an ideal world, every project in the portfolio would have dedicated screenshots from specific tools and processes. But preparing hundreds of such screenshots is a Herculean task. I went with the approach: **Done is better than perfect.** Instead of wasting time taking screenshots or hunting for stock photos, I went with generative graphics. * I used **Nano Banana MCP**. * I provide the content, the agent generates the graphic, and a script (converter) automatically converts PNG to WebP. This way, the site is visually consistent, maintains a tech/cyber aesthetic, and I don't have to worry about "gaps" in the content. ![og-zapier-vs-make-vs-n8n-wybor-narzedzia](/images/og-zapier-vs-make-vs-n8n-wybor-narzedzia.webp) ## 4. Reality Check: Experience vs. New Frameworks I have 15 years of IT experience (.NET, Python, JS), but frameworks like React and Vite were new to me. I had to learn them. And here's the key point: **Programming knowledge is essential.** Thanks to my experience in system design, the code generated by AI is understandable to me. I can assess its correctness before it hits production. Without that, I would have drowned in errors. * Agents (even Claude Sonnet 4.5 or Gemini 3 Pro) can get stuck in loops. * There are hallucinations of non-existent libraries. If I didn't understand the fundamentals, I wouldn't have been able to "unwind" the errors AI introduced in more complex logic. ### Code Pedantry I'm pedantic about keeping files organized (Clean Code). Clear folder structure and separation of concerns are sacred to me. AI tends to dump everything into one bucket. My role was to enforce that structure. Thanks to this, the project is easy to maintain and reorganize, rather than being "spaghetti code" spit out by a machine. ## 5. Insurance Policy: Tests and Code Review Agent In a hobby project, chaos comes easily. To prevent that, I implemented two levels of safeguards: 1. **Comprehensive E2E tests (Playwright):** Every change is verified by automated tests. I'm confident that a new feature hasn't "broken" an old one. 2. **Code Review Agent in Cursor.sh:** This is a brilliant feature. The agent analyzes changes in the last commit *before* pushing to the repository. It caught quite a few logical errors and potential issues that I might have overlooked. ![playwright_report](/images/playwright_report.webp) ## 6. Costs: How Much Does a "Free" Programmer Cost? This is an interesting comparison. Throughout the entire project, I consumed about **60 million tokens**. The input/output split is roughly 80/20. If I had paid API rates (e.g., Claude Sonnet 4.5 -- $3 input / $15 output), the cost would have been about **$325 USD**. Actual cost? * PRO+ plan in Cursor.sh: **$60 USD**. * Antigravity (included in Google Workspace): **$0 USD** (included in the company package). The savings are enormous. Of course, there are limits -- during intense sessions, I'd occasionally see a message about exceeding them. I'd simply switch models (jumping between Gemini 3 Pro High and Claude Sonnet 4.5). Limits refresh every few hours, so with regular work it's not a blocker. ![cursor_usage](/images/cursor_usage.webp) ## Summary The [pawel.lipowczan.pl](https://pawel.lipowczan.pl) project is proof that in 2025, the programmer's role is evolving. We're no longer syntax craftspeople -- we're becoming architects managing a team of digital agents. You might not be a designer. You might not know the latest framework inside and out. But if you have an engineering mindset, a commitment to quality (and tests!), and the ability to orchestrate AI -- you can build things that previously would have required an entire team. Feel free to check out the results and do a code review! Feedback is always welcome.

Need support developing a product with AI?

I can help you choose the right technology, design the architecture, and implement best practices. From MVP through scaling to optimizing development processes with AI.

Book a free consultation
## FAQ
### Can AI replace programmers in 2025? AI is a junior developer on steroids -- it speeds up work 10x but requires supervision from an experienced engineer. Agents can get stuck in loops, hallucinate non-existent libraries, and generate "spaghetti code" without imposed structure. Engineering fundamentals are essential for assessing code correctness and unwinding errors.
### What skills do programmers need to work effectively with AI? Knowledge of system architecture, clean code, and separation of concerns -- AI tends to dump everything into one bucket. The ability to read and assess code, even in an unfamiliar framework. Commitment to quality: E2E tests and code review before every commit. Without these fundamentals, you'll drown in errors.
### How much does it cost to develop a project with AI compared to traditional programming? A PRO+ plan in Cursor.sh is about $60/month with intensive work. For comparison: the same 60 million tokens via API would cost ~$325. The savings are enormous, but there are limits -- during intense sessions you may need to switch models or wait for renewal. Regular work can proceed without blockers.
### How do you ensure the quality of AI-generated code? Two levels of safeguards: E2E tests (Playwright) that automatically verify every change, and a code review agent that analyzes changes before commit. The agent catches logical errors and potential issues that are easy to overlook. Without tests and review, chaos in the project is inevitable.
### Can a person without programming experience build a professional application with AI? No -- a professional, secure, and scalable service requires technical knowledge. "Vibe coding" and chatting with a chatbot produce visual results but not full control over the code. AI is excellent at supporting experienced engineers but doesn't replace programming fundamentals for complex projects.
--- # Every Company Operates Suboptimally - How to Stop Lying to Your Employees and Start Fixing Processes Source: https://pawel.lipowczan.pl/en/blog/every-company-operates-suboptimally Published: 2025-12-01 Do you have that person in your company? The rock. Someone who never lets you down, even when everything around them is falling apart. And for the third time this month, they come to you with the same absurd problem -- a system bug that's blocking their work. You look at the screen, then at them, and feel that burning sting of shame. You say: _"Don't worry, I'll take care of it"_, but deep down you know you're lying. Not out of bad faith. You're lying because the technology that was supposed to help is making you a liar in the eyes of your best people. You're stuck with rigid systems where every change is a Moon-landing-scale project. If this scenario sounds familiar, you're not alone. At Automation House, we've mapped over 400 processes and the conclusion is clear: **every company operates suboptimally**. The only question is: how quickly can you find those spots and fix them? ## Without a Map, There's No Navigation Imagine you want to reach a destination in unfamiliar terrain. Without a map, you wander. It's the same in business. A process map isn't just documentation -- it's a navigation tool for three groups: 1. **Business:** Gains an understanding of how the company _actually_ works (management's assumptions often miss the mark). 2. **Users:** Receive clear instructions and faster onboarding. 3. **IT/Implementers:** Can precisely design architecture, transfer project knowledge, and find bottlenecks. Proof? Research shows that process mapping in healthcare alone has reduced patient wait times by **20-45%**. If it works in an environment as complex as a hospital, it'll work in your company too. ## Why Most Process Maps Are Useless Many managers try to map processes but do it poorly, choosing the wrong tools: - **SIPOC (tables):** Great for analysts, incomprehensible for business people. Hard to draw conclusions from at a glance. - **BPMN (Business Process Model and Notation):** The corporate standard, but overly complex. Too many logic gates and symbols make the map unreadable for the average employee. - **Basic Flowchart:** Too simple. Shows "what" happens but often misses "who" does it and "with what." ### The Golden Mean: Extended Flowchart At Automation House, we developed our own method. Our process map must include four key elements for each step: 1. **Action:** What happens? 2. **Actor:** Who does it? 3. **Tool:** What do they use? (e.g., Excel, CRM, Slack) 4. **Mode:** Manual or Automated? This way, you can immediately see where a person is doing a robot's job (copy-paste) and where integrations between systems are missing. ## How to Find the "Broken Pipes" Once you have your current-state map (AS-IS), finding optimizations becomes straightforward. Look for places where: - The most errors occur. - The process takes the longest. - Data is being manually re-entered (error risk, time waste). - The impact on the team will be greatest. **Elon Musk's golden rule:** Before you automate anything, ask yourself: _Does this step even need to exist?_ > "Probably the worst thing you can do is optimize something that shouldn't exist in the process in the first place." First remove, then simplify, and only then automate. ## Case Study: El Padre - How to Speed Up Proposals by 50% Theory is theory, but let's look at practice. Event agency **El Padre** came to us with a problem: creating proposals (especially smaller ones) was too time-consuming and not cost-effective. Knowledge from previous projects was scattered in employees' heads -- there was no central knowledge base. **What did we do?** 1. **Step 1: The Process "Ear" (Fireflies.ai)** We deployed an AI tool that records client meetings and creates transcripts. No more manual note-taking and losing details. 2. **Step 2: The Central Brain (Airtable)** We created a knowledge base where transcripts, cost estimates, and project data all converge. 3. **Step 3: Automation (Make & AION)** We built "AI assistants" (powered by our AION tool) that handle specific tasks: - **Briefing:** Generates a brief based on the meeting transcript. - **Event Ideas:** Suggests event ideas based on the brief and agency history. - **Financial Planner:** Creates a cost estimate based on the brief and agency history. - **Offer Generator:** Produces a proposal from previously generated cost estimates, ideas, and the brief. **Results:** - **10-50%** faster proposal preparation (biggest gains on smaller projects). - **10-15%** production department productivity increase. - **30 people** in the company genuinely supported by AI in their daily work. ## Summary Technology isn't meant to complicate life -- it's meant to build Operational Excellence. You don't need to deploy complex ERP systems right away. Start with a map. Find where your company is "bleeding" time and people's energy. If you want to stop lying to your employees that "it'll work out somehow," start by mapping one process this week.

Want to map and optimize your company's processes?

I can help you find process bottlenecks, identify automation opportunities, and deploy solutions that save your employees time. From mapping through analysis to implementation.

Book a free consultation
## FAQ
### Why does every company operate suboptimally? Processes grow organically over the years, management's assumptions often miss reality, and employees do robot work (copy-pasting between systems). From mapping over 400 processes, one conclusion emerges: suboptimal spots are everywhere. The only question is how quickly you find and fix them.
### What elements should an effective process map include? Four elements for each step: action (what happens), actor (who does it), tool (with what -- Excel, CRM, Slack), and mode (manual or automated). This immediately shows where a person is doing a robot's job and where integrations between systems are missing.
### How do you find optimization opportunities in business processes? Look for places where: the most errors occur, the process takes the longest, data is being manually re-entered (error risk, time waste), and where the change will have the greatest impact on the team. Before automating, ask: does this step even need to exist?
### Why are most process maps useless? SIPOC (tables) is incomprehensible for business people, BPMN is overly complex (too many gates and symbols), and a basic flowchart is too simple -- it shows "what" without "who" and "with what." The golden mean is an extended flowchart with four elements: action, actor, tool, mode.
### What is the right sequence of actions when optimizing processes? First remove (does this step need to exist?), then simplify (can it be shortened?), and only then automate. The worst thing you can do is automate something that shouldn't exist in the process in the first place. That's the golden rule before any optimization project.
--- _This article is based on Pawel Lipowczan's presentation "Every Company Operates Suboptimally" delivered at InfoShare Katowice 2025._ --- # Zapier vs Make vs n8n - how to choose the right automation tool for your team? Source: https://pawel.lipowczan.pl/en/blog/zapier-vs-make-vs-n8n-tool-choice Published: 2025-11-17 # Zapier vs Make vs n8n - how to choose the right automation tool for your team? Choosing the wrong automation tool isn't just wasted money - it's **months of wasted time**, hundreds of rewritten workflows, and thousands of dollars on migration when you finally decide to switch. I've seen it dozens of times: teams choose a platform based on a feature list, then get stuck because nobody knows how to use it. After implementing automation for over 100 clients at **Automation House**, I can say one thing: **there's no universal answer**. But there are specific criteria that determine whether a given tool will work for your team. In this article I'll show you a **decision framework** that will help you choose between **Zapier**, **Make**, and **n8n** based on what really matters: your team's competencies, operational scale, budget, and security requirements. ## Why it's not just about features All three platforms do the same thing - **connect your apps without coding**. But the devil is in the details: - **Zapier** has 6,000+ integrations and is so simple your mom could use it - **Make** (formerly Integromat) gives you a visual canvas where you see the entire workflow logic - **n8n** is a technical team's dream: open-source, self-hosted, unlimited possibilities Most companies I work with **waste 15-25 hours per week** on repetitive tasks: data entry, notifications, status updates, cross-platform syncing. Automation kills that time waste. But when they choose a tool based on a feature list instead of team capabilities, they hit a wall and have to rebuild everything from scratch. ## Zapier - for non-technical teams that need results now ### Who is it for? Zapier is the **"it just works"** platform. Simple trigger-action configuration that non-technical teams can deploy in minutes. **Best for:** - RevOps departments connecting HubSpot/Salesforce/Slack - Marketing automation and lead routing - IT workflows (onboarding, ticketing) - Startups without a technical CTO ### Pros ✅ **6,000+ integrations** - if an app exists, Zapier supports it ✅ **Zero learning curve** - non-technical teams start in 5 minutes ✅ **Excellent documentation** - template library, community, video guides ✅ **Instant results** - first workflow in 10 minutes ✅ **Largest community** - every problem has a solution on the forum ### Cons ❌ **Task limits skyrocket** - 5-step Zap x 100 runs = **500 tasks** ❌ **Pricing escalates fast** - the free 100 tasks vanish in a blink ❌ **Limited conditional logic** - hard to build complex decisions ❌ **Debugging is a nightmare** - when something breaks, it's hard to find the cause ❌ **Vendor lock-in** - migrating to another platform = rewriting from scratch ### Example use cases **Lead routing:** ```text Contact form → Zapier → - Add lead to HubSpot - Send Slack notification - Create task in Asana - Send welcome email ``` **Employee onboarding:** ```text New record in BambooHR → Zapier → - Create Google Workspace account - Add to Slack channels - Send welcome email with checklist - Create tasks for manager ``` ### Pricing - **Free:** 100 tasks/month - **Starter:** $19.99 (750 tasks) - **Professional:** $49 (2,000 tasks) - **Team:** $299 (50,000 tasks) **Warning:** Every step in a Zap is a separate task! A 5-step Zap run 100 times = 500 tasks. ### When to choose Zapier? ✅ Your team is non-technical (marketing, sales, ops) ✅ You need results immediately, no training needed ✅ You're connecting niche apps (they have the most integrations) ✅ You're running pilots and proofs of concept ✅ Scaling isn't your priority (< 5,000 tasks/month) ## Make - for visual thinkers with ambition ### Who is it for? Make is the platform for teams that **think visually** and need more power than Zapier but without the technical complexity of n8n. **Best for:** - Creative agencies with content pipelines - Marketing automation with personalization - Teams with power users - Processes requiring complex "if-this-then-that" logic ### Pros ✅ **Visual workflow builder** - you see the entire process on canvas ✅ **Better value for money** - 10x more operations for the same price ✅ **Advanced logic** - routers, filters, iterators, error handlers ✅ **Transparent debugging** - each step shows input/output data ✅ **Operations ≠ steps** - each step is 1 operation (doesn't multiply like in Zapier) ### Cons ❌ **Steeper learning curve** - the visual builder takes getting used to ❌ **Fewer integrations** - 1,800+ apps (vs 6,000+ in Zapier) ❌ **Uneven documentation** - some modules are poorly documented ❌ **Interface can overwhelm** - initially chaotic on canvas ### Example use cases **Content pipeline with categorization:** ```text Webhook → Make → ├─ If type = "blog post" │ └─ Add to WordPress + notify writers ├─ If type = "social media" │ └─ Schedule in Buffer + notify social team └─ If type = "newsletter" └─ Add to Mailchimp + notify subscribers ``` **Marketing campaign with personalization:** ```text New subscriber → Make → ├─ Fetch data from CRM ├─ Router by segment: │ ├─ B2B → Email sequence A │ ├─ B2C → Email sequence B │ └─ Enterprise → Notify sales team └─ Add to appropriate remarketing list ``` ### Pricing - **Free:** 1,000 operations/month - **Core:** $9 (10,000 operations) - **Pro:** $16 (10,000 operations + premium apps) - **Teams:** $29 (10,000 operations + team features) **Key difference:** In Make, each step = 1 operation (no multiplication!). A 10-step workflow x 1,000 runs = 10,000 operations. ### When to choose Make? ✅ You want more power than Zapier without n8n's technical complexity ✅ Your team thinks visually and likes to "see" the logic ✅ Workflows have many branches and conditions ✅ You're building for clients and need to show the logic ✅ You're looking for the best value for money ## n8n - for technical teams with requirements ### Who is it for? n8n is an **open-source powerhouse** for technical teams that want full control over automation. **Best for:** - Teams with developers/DevOps - Regulated industries (healthcare, finance, legal) - High-volume automation (50K+ tasks/month) - Custom API integration needs - Agencies building automation products for clients ### Pros ✅ **Open-source (MIT license)** - full access to source code ✅ **Self-hosted = zero subscription costs** - you only pay for infrastructure ✅ **Unlimited workflow steps** - no step limits ✅ **Custom code nodes** - JavaScript in every step ✅ **Full data control** - for compliance (HIPAA, GDPR, SOC2) ✅ **API-first approach** - easy integration with custom systems ✅ **Great for AI agents** - advanced workflows with LLMs ### Cons ❌ **Requires DevOps skills** - Docker, databases, SSL, backups, monitoring ❌ **Self-hosting = maintenance overhead** - updates, security patches ❌ **Smaller community** - fewer templates and examples ❌ **Cloud hosting more expensive** - than Make (if you don't self-host) ❌ **Security is your responsibility** - you handle security yourself ### Example use cases **Healthcare data pipeline (HIPAA compliant):** ```text Patient intake form → n8n (self-hosted) → ├─ Encrypt PHI data ├─ Store in compliant database ├─ Notify medical staff (secure channel) └─ Log audit trail ``` **AI agent workflow:** ```text User query → n8n → ├─ Pre-process with custom code ├─ Route to appropriate LLM (OpenAI/Claude/Local) ├─ Post-process response ├─ Store in vector database └─ Return formatted result ``` **Multi-tenant automation product:** ```text Client webhook → n8n → ├─ Identify tenant ├─ Load tenant-specific config ├─ Execute custom workflow ├─ Bill based on usage └─ Store metrics per tenant ``` ### Pricing **Self-hosted:** - **Software:** $0 (MIT license) - **Infrastructure:** $10-50/month (VPS: DigitalOcean, Hetzner, AWS) - **DevOps time:** 5-10h/month (setup, maintenance) **n8n Cloud:** - **Starter:** $20 (2,500 workflow executions) - **Pro:** $50 (10,000 executions) - **Enterprise:** Custom pricing ### When to choose n8n? ✅ You have a developer or DevOps person on the team ✅ Data privacy is critical (healthcare, finance, legal) ✅ You're scaling above 50K tasks/month ✅ You need custom code or non-standard APIs ✅ You're building automation products for multiple clients (multi-tenant) ✅ Compliance requirements (HIPAA, GDPR on-premises) ## Code with AI agents - an option nobody considers (but should) ### Who is it for? 2026 changed the game. AI agents like Claude Code, Cursor, and GitHub Copilot did to writing code what Zapier did to integrations: **lowered the barrier to entry to zero**. **Best for:** - Anyone who has a developer (even a junior) with access to an AI agent - Companies that want full control without vendor lock-in - Use cases requiring non-standard logic and many integrations - Anyone building something specific to their business (micro-tools) ### How does it work in 2026? The old calculation: custom code = weeks of work, high cost, hard to maintain. **New calculation**: got API documentation? Hand it to the AI agent. Integration ready in hours, not weeks. Example: you need a webhook that takes data from Notion, enriches it via an external API, filters by 5 criteria, and saves to Airtable + sends a Slack notification? In Make: 45 minutes of configuration. In code with Claude Code: 2 hours and you have a deployment. But the code is yours. No run limits. You're not paying $50/month forever. ### Pros ✅ **Zero vendor lock-in** - code runs anywhere, migrate wherever you want ✅ **Any API without waiting** - give the agent documentation, it has the integration in hours ✅ **No limits** - zero artificial limits on steps, runs, data ✅ **Cheapest at scale** - hosting $5-20/month vs hundreds in subscriptions ✅ **Full debuggability** - no black boxes, every step in the logs ✅ **Composable** - like Lego bricks: each micro-tool does one thing and does it well ✅ **Best flexibility** - changing logic means changing code, not fighting with a platform UI ### Cons ❌ **Requires a developer** - junior + AI agent is the minimum, but it has to be someone with a tech background ❌ **Maintenance** - code needs maintaining (though AI helps with that too) ❌ **Setup time** - first deployment slower than "click-click in Zapier" ❌ **Infrastructure** - hosting, deployment, monitoring (but tools like Railway/Fly.io minimize overhead) ### Example: micro-tool instead of a platform **Task**: Process 500 leads per day from 3 sources, deduplicate, enrich with Clearbit data, save to CRM, and notify sales. **In Zapier**: $300+/month, 5-step Zap x 500 = 2,500 tasks/day **In Make**: $200+/month, complex scenario **In code + Claude Code**: 1-2 days to build, $10/month on infrastructure, unlimited ```python # What an AI agent will write for you in hours: # - Webhook accepting leads from 3 sources # - Deduplication logic # - Clearbit enrichment # - CRM write # - Slack notification # Zero platform. Zero limits. Zero vendor lock-in. ``` ### Pricing - **Infrastructure**: $5-20/month (Railway, Fly.io, Render) - **Developer time**: 2-8h one-time (with AI agent instead of 2-4 weeks) - **Maintenance**: 1-2h/month (with AI help) - **Total**: ~$20/month + one-time cost ### When to choose Code with AI agents? ✅ You have a developer (junior/mid) with access to Claude Code / Cursor ✅ You want zero vendor lock-in ✅ The logic is specific to your business ✅ Scale >20K operations/month (where subscriptions hurt) ✅ You want to build micro-tools, not deploy a monolithic platform ✅ Integrations with APIs that don't have official connectors in Zapier/Make ## Decision framework - how to actually choose? Instead of guessing, use this simple framework: ### Question 1: What competencies does your team have? - **Non-technical** (marketing, sales, operations) → **Zapier** - **Power users**, visual thinkers → **Make** - **Developers**, DevOps, technical team → **n8n** ### Question 2: What's the scale of operations? - **< 5,000 tasks/month** → **Zapier** or **Make** - **5,000 - 50,000 tasks/month** → **Make** - **> 50,000 tasks/month** → **n8n** (self-hosted) ### Question 3: What's the budget? - **Minimal budget, quick wins** → **Zapier Free/Starter** - **Best value for money** → **Make** - **Long-term, high-volume** → **n8n self-hosted** ### Question 4: What are the security requirements? - **Standard SaaS security** → **Zapier** / **Make** - **Data residency, compliance** → **n8n self-hosted** - **HIPAA, GDPR, SOC2 on-premises** → **n8n self-hosted** ### Question 5: How complex are the processes? - **Simple trigger-action** (A → B → C) → **Zapier** - **Multi-step with conditions** (if-else, routers) → **Make** - **Complex logic + custom code** → **n8n** ### Question 6: Do you have a developer with access to an AI agent? - **Yes** → consider Code with AI as the first option (zero vendor lock-in, lowest cost at scale) - **No** → go back to questions 1-5 (Zapier/Make/n8n) ## Common mistakes when choosing (and how to avoid them) ### Mistake 1: Choosing n8n without technical resources **What happens:** - Deploy on VPS, everything works - After a week: SSL certificate issue - After a month: database full, backup not working - After 3 months: security vulnerability, no updates **Solution:** ✅ Hire a DevOps consultant (5-10h/month) ✅ Use managed n8n hosting (much more expensive, but no headaches) ✅ Or... choose Make instead of n8n ### Mistake 2: Starting on Zapier and hitting the task limit **What happens:** - Start with Zapier Free (100 tasks) - After a week: upgrade to Starter ($20, 750 tasks) - After a month: upgrade to Professional ($50, 2,000 tasks) - After a quarter: $300/month, and the workflows are simple **Why?** 5-step Zap x 100 runs = 500 tasks! **Solution:** ✅ If you see yourself scaling beyond 5K tasks, **go straight to Make** ✅ Prototype in Zapier, production in Make ✅ Zapier → Make migration means rewriting from scratch (plan ahead) ### Mistake 3: Choosing based on feature list instead of team fit **What happens:** - CTO picks n8n because "it's open-source and has all the features" - The marketing team can't use it - Developers don't have time to build workflows - Result: 0 deployed automations after 3 months **Solution:** ✅ **The best tool is the one your team will actually use** ✅ Simplicity > functionality (if nobody can use it) ✅ Start with a pilot project with the team that will use the tool ### Mistake 4: Not accounting for Total Cost of Ownership (TCO) **Zapier TCO:** - Subscription: $50-300/month - Team time: 2h/month (maintenance) - **Total: $50-300/month** **Make TCO:** - Subscription: $16-50/month - Learning curve: 10h (one-time) - Team time: 3h/month (maintenance) - **Total: $16-50/month** **n8n TCO (self-hosted):** - Infrastructure: $20-50/month - DevOps time: 10h/month x $50/h = $500 - **Total: $520-550/month** **n8n TCO (cloud):** - Subscription: $50-200/month - Team time: 3h/month - **Total: $50-200/month** **Code + AI agent TCO:** - Infrastructure: $10-20/month (Railway, Fly.io, Render) - Dev time: 2-8h one-time (with AI agent) - Maintenance: 1-2h/month - **Total: ~$20/month + one-time build cost** | Tool | Monthly cost | One-time cost | |--------------------|------------------|-------------------| | Zapier | $50-300 | minimal | | Make | $16-50 | low | | n8n self-hosted | $520-550 | high | | Code + AI agent | $10-20 | medium (1x) | **Takeaway:** n8n self-hosted only makes sense for **high-volume** (>50K tasks) or **compliance requirements**. Code + AI agent wins when you have a developer and want full control without growing subscriptions. ## Migration strategies between platforms ### From Zapier to Make **When?** Zapier costs > $100/month, and workflows are of medium complexity. **How?** 1. Identify the simplest Zaps (3-5 steps) 2. Rewrite them in Make (visual canvas helps with optimization) 3. Test in parallel for a week 4. Only disable Zaps after verification 5. Gradually migrate more complex workflows **Time required:** 2-4h per workflow ### From Make to n8n **When?** Make costs > $200/month or compliance requirements. **How?** 1. Deploy n8n on managed hosting (Railway, Render) 2. Export workflows from Make as JSON (partially compatible) 3. Migrate non-critical workflows first 4. Test thoroughly (differences in nodes) 5. Gradual migration of production workflows **Time required:** 5-10h per workflow (nearly complete rewrite) ### Multi-platform approach You don't have to choose just one platform! **Strategy:** - **Zapier** - quick wins, prototypes, proofs of concept - **Make** - production workflows, team standards - **n8n** - high-volume, sensitive data, complex logic **Example:** - Marketing uses Zapier (lead routing, simple integrations) - Product team uses Make (onboarding, notifications) - Engineering team uses n8n (data pipelines, AI agents) ## Case studies - real world scenarios ### Case Study 1: A marketing startup chose Zapier **Team:** 3 people (CEO, marketer, designer) **Problem:** Manual lead routing from 5 sources **Solution:** 3 simple Zaps **Workflow:** ```text 1. Form → HubSpot + Slack 2. LinkedIn Lead Gen → HubSpot + Email 3. Chatbot → HubSpot + Asana task ``` **Results:** - ✅ Deployment: 2 hours - ✅ ROI: day one (saved 5h/week) - ✅ Cost: $50/month (Professional plan) - ✅ Satisfaction: 10/10 **Why Zapier?** Non-technical team, simple workflows, instant results. ### Case Study 2: A creative agency chose Make **Team:** 15 people (designers, copywriters, project managers) **Problem:** Content chaos - 5 content sources, 10 publication channels **Solution:** 25 complex workflows with categorization **Workflow (example):** ```text Content submission → Make → ├─ Classify content type (AI) ├─ Router by type: │ ├─ Blog → WordPress + notify writers │ ├─ Social → Buffer (multi-channel) + notify social │ ├─ Newsletter → Mailchimp + notify subscribers │ └─ Client → Dropbox + notify client success ├─ Update project status (Asana) └─ Log metrics (Google Sheets) ``` **Results:** - ✅ Saved: 20h/week - ✅ Cost: $150/month (vs $800 on Zapier) - ✅ Workflows: 25 active, averaging 12 steps each - ✅ Complexity: impossible to achieve in Zapier **Why Make?** Power users, visual logic, best value for money. ### Case Study 3: A software house chose n8n **Team:** 30 people (15 developers, 10 product, 5 ops) **Problem:** 50K+ tasks/month, GDPR requirements **Solution:** n8n self-hosted on AWS **Workflows (examples):** ```text 1. User registration → Encrypt PII → Store EU database → Email 2. Payment webhook → Process → Update CRM → Generate invoice 3. Support ticket → Classify (AI) → Route → Notify → Track SLA 4. CI/CD webhook → Test → Deploy → Notify → Update docs ``` **Results:** - ✅ Volume: 50,000+ workflow executions/month - ✅ Cost: $30/month (AWS EC2 t3.medium) - ✅ vs Make: $400+/month (at this volume) - ✅ vs Zapier: $1,200+/month - ✅ Compliance: GDPR-compliant (EU-hosted) **Why n8n?** Technical team, high-volume, compliance requirements, ROI after 2 months. ### Case Study 4: A software startup chose Code + AI **Team:** 2 developers + Claude Code **Problem:** 30K leads/month from 5 sources, complex deduplication and data enrichment logic **Solution:** Python micro-service + webhooks, deployed on Railway **Workflow:** ```text Webhook (5 sources) → ├─ Deduplication logic (custom) ├─ Clearbit enrichment ├─ Filtering by business criteria ├─ Save to CRM └─ Slack notification ``` **Results:** - ✅ Build time: 3 days (vs estimated 3 weeks without AI) - ✅ Cost: $15/month (vs $400 on Make at this scale) - ✅ Vendor lock-in: $0 (zero) - ✅ Custom logic: unlimited, change = change a line of code **Why Code?** Developer + Claude Code = no-code agility + code power. At 30K leads/month Make would cost $200-400/month. Code: $15/month and unlimited operations. ## The future of no-code automation ### Trends I'm observing **1. AI-driven automation** - ChatGPT/Claude nodes in every platform - Intelligent categorization and routing - Content generation within workflows **2. Conversational workflow creation** - "Create a workflow that does X" → done - Citizen developers vs technical teams - Democratization of automation **3. Platform consolidation** - All-in-one (automation + data + AI) - Kestra, Temporal, Prefect - the new generation **4. Regulatory compliance automation** - GDPR, HIPAA, SOC2 out-of-the-box - Automated audit trails - Self-hosted renaissance ### My recommendations for 2026 **For startups:** - Start simple: **Zapier** - Scale smart: **Make** when you exceed 5K tasks - Have a developer + AI: **Code** right away - zero vendor lock-in - Go technical: **n8n** if compliance or high-volume **For agencies:** - Default choice: **Make** (best value) - Client work: Zapier for simple, Make for complex - Product building: **n8n** for multi-tenant SaaS - Internal tools: **Code + AI** for custom micro-tools **For enterprise:** - Departmental: **Zapier**/Make for individual departments - Central automation: **n8n** self-hosted for IT - Tech team + AI: **Code** for specific tools and integrations - Governance: Multi-platform approach with central oversight ## Summary - Quick Decision Guide ### Want the simplest start? → **Zapier** - Non-technical team - < 5,000 tasks/month - Simple integrations - Results in 10 minutes ### Want the best value? → **Make** - Power users on the team - 5,000 - 50,000 tasks/month - Complex conditional logic - 10x more for the same money ### Want maximum control? → **n8n** - Technical team - > 50,000 tasks/month - Compliance requirements - Custom integrations ### Want maximum flexibility and zero limits? → **Code + AI agent** - Have a developer with access to Claude Code / Cursor - Specific business logic - Scale > 20K operations/month - Zero vendor lock-in, zero artificial limits ### Not sure? → **Multi-platform approach** - Zapier for prototypes - Make for production - n8n for specific use cases **Remember:** The best tool is the one your team will actually use. Match the platform to your team's competencies, not the other way around.

Need help choosing the right automation tool?

I'll help you analyze your company's needs, choose the right platform (Zapier, Make, or n8n), and implement the first workflows. From process audit through tool selection to team training.

Book a free consultation
## FAQ
### How to choose between Zapier, Make, and n8n for your team? The choice depends on three factors: team competencies, operational scale, and budget. Zapier for non-technical teams (<5K tasks/month), Make for power users looking for the best value (5-50K tasks), n8n for technical teams with compliance requirements (>50K tasks).
### Which automation platform offers the best value for money? n8n offers the best value for complex workflows - 1 credit is the execution of an entire workflow regardless of the number of modules. In Make, each module consumes a separate credit. For simple workflows (3-5 steps) Make may be more cost-effective due to a lower per-credit price, but for complex processes n8n wins economically.
### Is self-hosted n8n really cheaper than Zapier and Make? Only at very large scale (>50K tasks/month) or with compliance requirements. Infrastructure costs are $20-50/month, but add 10h/month of DevOps work ($500+). Make at $50/month handles most cases without maintenance overhead.
### When is it worth using multiple automation platforms simultaneously instead of one? A multi-platform approach works when different departments have different needs. A typical strategy: Zapier for prototypes and quick tests, Make for production workflows, n8n for high-volume or sensitive data. This eliminates the compromises that come from choosing a single tool.
### Why does migration from Zapier to Make require rewriting workflows from scratch? The platforms use different data models and workflow structures - there's no direct compatibility. Zapier counts every step as a separate task, Make treats the entire workflow as one operation. Migration requires 2-4h per workflow, but pays off long-term through 5-10x lower operational costs.
--- # How Event Agency El Padre Accelerated Proposal Creation by Up to 50% with AI Source: https://pawel.lipowczan.pl/en/blog/el-padre-ai-offer-automation Published: 2025-11-16 # How Event Agency El Padre Accelerated Proposal Creation by Up to 50% with AI In the competitive world of events, where every minute counts, event agency **El Padre** had to face a challenge that many companies in the industry deal with: how to increase the number of proposals submitted without sacrificing their high quality? The answer turned out to be deploying the **AION** platform with advanced AI support. The results? **10-50% faster proposal preparation**, **10-15% higher production department productivity**, and **25-30 people supported by AI** in their daily work. In this case study, I'll walk you through step by step how we carried out the implementation and what concrete business benefits this solution delivered. ## The Business Problem: When Quality Meets Time Pressure El Padre is a renowned event agency that prepares comprehensive proposals for its clients -- from creative concepts through visualizations to detailed cost estimates and schedules. Every proposal is a custom project requiring involvement from two key departments: ### Time Is Money -- Literally Before deploying AION, preparing a single proposal took: - **Creative department:** 6-10 hours (research, concept, visualizations) - **Production department:** 4-6 hours (pricing, cost estimation, logistics) - **Total:** up to **16 hours per proposal** For smaller projects (50-100K PLN), this workload was **unprofitable** -- but without a professional proposal, winning the bid was nearly impossible. ### Key Challenges 1. **Overloaded creative team** -- too many projects, too little time for each 2. **Time-consuming creative and production work** -- manual cost estimation, searching through previous proposals, creating visualizations 3. **No knowledge centralization** -- meeting transcripts, cost estimates, and archived proposals were scattered across various OneDrive folders 4. **Disproportionate proposal preparation cost** -- especially for smaller projects where the margin didn't justify the effort The agency faced a dilemma: either grow the team (which generates costs) or find a way to speed up processes without sacrificing quality. ## The Solution: AION as the Central Hub for Proposal Processes We decided to deploy the **AION** platform -- a system for managing AI-supported processes. The key to success wasn't just introducing AI tools, but above all **integrating and centralizing data** and **adapting the workflow to the team's real needs**. ### Technology Stack - **AION** -- platform for managing AI-supported processes - **OneDrive** -- integration with existing file system (automatic synchronization) - **AI transcription tools** -- automatic recording and processing of meetings - **AI assistants** -- targeted support for specific proposal stages - **Knowledge base search system** -- quickly finding information from previous projects ## Implementation: How We Did It in 6 Weeks We split the implementation into three main phases, completed over approximately **6 weeks**: ### Step 1: Data Integration and Centralization (Weeks 1-2) The first step was organizing the information chaos. We automated file synchronization from OneDrive to the AION platform, creating one central place for: - **Client meeting transcripts** -- every conversation automatically recorded and processed - **Knowledge about all projects** -- collaboration history, preferences, notes - **Cost estimates and pricing** -- price database and calculation templates - **Archived proposals** -- as a pattern base and reference points for new projects This gave the team **instant access** to all organizational knowledge -- without searching through dozens of folders. ### Step 2: Deploying Intelligent Tools (Weeks 3-4) In the second phase, we launched AI tools to support specific tasks: - **Automatic meeting recording and transcription** -- every client meeting was recorded (with consent) and then processed into a transcript and structured data - **Processing transcripts into structured data** -- AI extracted key information: budget, preferences, requirements, deadlines - **Visualization generator** -- AI helped quickly create initial visual concepts for event proposals - **Knowledge base search system** -- ability to filter by meetings, projects, clients, and quickly find similar cases ### Step 3: Workflow Implementation and Support (Weeks 5-6) The final phase was adapting the system to daily work and training the team: - **Preparing targeted AI assistants** -- each stage of the proposal process (research, concept, pricing, finalization) received a dedicated AI assistant - **Technical support during implementation** -- daily standups, resolving issues in real time - **Team training** -- **25-30 people** went through workshops and onboarding with AION - **Pilot tests and optimization** -- fine-tuning the workflow based on team feedback ## Results: The Numbers Speak for Themselves After 6 weeks of implementation, El Padre started seeing the first effects, which only grew over time: ### Key Metrics - **10-50% faster proposal preparation** -- depending on project complexity (simple proposals up to 50% faster, complex ones ~10-20%) - **10-15% higher production department productivity** -- thanks to automated pricing and cost estimation - **25-30 people supported by AI** -- practically the entire team uses AION in their daily work - **Dramatic increase in monthly proposals submitted** -- without growing the team ### ROI: Time Savings The numbers are impressive: - **Average 5-8 hours saved per proposal** - At **15 proposals per month** = **75-120 hours saved** - That's the equivalent of **2-3 full-time employees** ### Business Benefits **Efficiency gains:** - The production department can handle more projects without additional hires - The creative team has more time for innovative concepts - Ability to bid on smaller projects (previously unprofitable) **Quality and consistency:** - All proposals maintain a high standard thanks to knowledge base access - Lower risk of pricing errors (automated cost estimation) - Faster access to archived projects as templates **Brand image and competitiveness:** - Strengthened image as an innovative, modern agency - Faster response time to proposal requests (competitive advantage) - Ability to serve more clients **Additional benefits:** - Knowledge centralization -- easier onboarding for new employees - Better project documentation -- everything in one place - Scalability -- the solution grows with the company ## Client's Voice: What Does El Padre Say? > _"The AION deployment significantly simplified and accelerated our event proposal preparation process in many aspects. We appreciate the flexibility of the solution and the support of the implementation team. Every day we discover new applications for AION in our organization and take full advantage of its capabilities, even though we're only using a fraction of its potential."_ > > **Jakub Cwiklinski** > Vice President at El Padre That last point is particularly interesting -- the El Padre team is using "just a fraction of the potential" of AION, which means **benefits will continue to grow over time** as they discover new applications for the platform. ## Key Takeaways: What Can You Learn from This Case Study? ### 1. AI in the Event Industry Is the Present, Not the Future El Padre shows that you can achieve **up to 50% process acceleration** -- and in an industry that seems highly creative and difficult to automate. ### 2. Data Integration and Centralization Is Key Without a unified knowledge base, AI has nothing to draw from. The first step -- organizing data -- was the foundation of success. ### 3. Implementation Doesn't Have to Be Long **6 weeks was enough** for 25-30 people to work more efficiently. You don't need months-long transformation projects. ### 4. ROI Is Measurable Time savings (75-120h monthly) translate directly into the ability to serve more clients without increasing costs. ### 5. AI Supports, It Doesn't Replace The creative team still creates unique concepts -- AI simply relieves them of tedious, repetitive tasks. ## Who Is This Solution For? AI implementation in proposal processes works particularly well for: - **Creative and event agencies** with a high volume of proposals - **Companies where the proposal process is complex and time-consuming** - **Organizations looking to scale** without proportionally growing their team - **Businesses seeking a competitive edge** through speed and innovation If your company faces similar challenges -- **overloaded teams**, **time-consuming processes**, **lack of knowledge centralization** -- it's worth considering AI deployment.

Want to automate your proposal processes?

I can help you build an intelligent system that speeds up proposal creation, leverages your team's knowledge, and reduces manual work. From process analysis through AI assistant design to deployment and training.

Book a free consultation
## FAQ
### What results can you achieve by automating the proposal process with AI? Typical results include 10-50% faster proposal preparation, 10-15% team productivity increase, and 75-120 hours saved monthly with 15 proposals. This allows you to serve more clients without growing the team and bid on smaller projects that were previously unprofitable.
### How long does it take to implement AI in the proposal process? A typical implementation takes about 6 weeks in three phases: data integration and centralization (weeks 1-2), AI tool deployment (weeks 3-4), workflow implementation and team training (weeks 5-6). You don't need months-long transformation projects to see the first results.
### How does AI support proposal creation in creative and event agencies? Automatic meeting transcription and extraction of key information (budget, requirements, deadlines), searching archived proposals as templates, automating cost estimates and pricing, and generating initial visualizations. AI handles the tedious tasks so the creative team can focus on unique concepts.
### How do you calculate ROI from proposal process automation? Measure average proposal preparation time before and after implementation, multiply the savings by the number of proposals per month and the team's hourly rate. Example: 5-8h savings x 15 proposals = 75-120h/month. Those hours can go toward more proposals or creative projects at no additional cost.
### What types of companies benefit most from AI-powered proposal automation? Creative and event agencies with a high volume of proposals, companies with complex and time-consuming proposal processes, and organizations looking to scale without proportionally growing their team. Key qualifying signals: overloaded teams, knowledge scattered across folders, and smaller projects that aren't worth the effort.
--- # Email Automation with AI - How Frontdesk AI Revolutionizes Customer Service Source: https://pawel.lipowczan.pl/en/blog/email-automation-frontdesk-ai Published: 2025-11-10 # Email Automation with AI - How Frontdesk AI Revolutionizes Customer Service Managing shared email inboxes is a challenge for many companies. Hundreds of messages per day, repetitive questions, the need for quick responses. Frontdesk AI solves this problem through intelligent automation. ## What is Frontdesk AI? Frontdesk AI is a system for automatically processing and categorizing incoming mail. It uses artificial intelligence (OpenAI) to analyze message content, classify it by category, and automatically generate responses. ## Key Features ### 1. Automatic Categorization The system analyzes the content of each message and automatically assigns it to the appropriate category: - Quote requests - Complaints - Technical questions - Invoices and payments - Spam ### 2. Intelligent Responses ```text For the most common questions, the system automatically generates responses based on a prepared FAQ knowledge base. ``` Each response is contextual and tailored to the specific customer question. ### 3. Routing to the Right People Messages requiring human intervention are automatically forwarded to the appropriate departments or individuals. ## Tech Stack - **Make** - workflow automation and integrations - **OpenAI GPT-4** - content analysis and response generation - **Gmail/Outlook API** - email integration - **Airtable** - knowledge base and message tracking ## Sample Workflow 1. A new email arrives at the shared inbox 2. A Make webhook captures the message 3. OpenAI analyzes the content and intent 4. The system checks the knowledge base in Airtable 5. If it finds an answer - it automatically sends a reply 6. If not - it forwards to the right person with context ## ROI and Savings A typical client saves: - **20-30 hours of work per month** on email handling - **90% reduction in response time** for standard questions - **100% availability** - the system works 24/7 - **Consistent quality** of responses ## Implementation The Frontdesk AI implementation process: 1. **Analysis** - mapping typical message categories 2. **Knowledge base** - preparing FAQ and response templates 3. **Configuration** - setting up workflows in Make 4. **Testing** - verifying operation on a test group 5. **Go-live** - launching for all incoming mail 6. **Optimization** - fine-tuning based on feedback Typical implementation time: 1-2 weeks. ## Summary Frontdesk AI is a solution for companies that: - Receive many repetitive questions - Need fast response times - Want to relieve their team from routine tasks - Care about customer service quality

Want to automate email handling in your company?

I'll help you implement intelligent email automation that relieves your team from routine questions and speeds up response times. From case analysis through configuration to testing and launch.

Book a free consultation
## FAQ
### What is Frontdesk AI and how does it automate email handling? Frontdesk AI is an automatic incoming mail processing system that uses OpenAI for content analysis. The system categorizes messages (inquiries, complaints, invoices, spam), automatically responds to standard questions from the FAQ knowledge base, and forwards complex issues to the right people with context.
### How does AI-powered automatic email categorization work? OpenAI GPT-4 analyzes the content and intent of each message, assigning it to a defined category. The system checks the knowledge base in Airtable - if it finds an answer, it automatically sends a reply. If not, it passes the message to the right person along with context and a suggested category.
### What savings does AI email automation deliver? A typical client saves 20-30 hours of work per month, reduces response time for standard questions by 90%, and gains round-the-clock availability (24/7). An additional benefit: consistent response quality - every customer receives the same standard of service regardless of the time of day.
### What types of companies benefit from an email automation system like Frontdesk AI? Companies that receive many repetitive questions, need fast response times, and want to relieve their team from routine tasks. It's ideal for customer service departments, info@ and support@ inboxes, where a large portion of inquiries concern FAQ, order status, or standard procedures.
### How long does it take to implement an email automation system? Typical implementation time is 1-2 weeks. The process includes: analyzing typical message categories, preparing the knowledge base and response templates, configuring workflows in Make, testing on a pilot group, launching for all mail, and optimization based on feedback.
--- # No-Code Lead Generation - How to Build a Lead Generation System Without Programming Source: https://pawel.lipowczan.pl/en/blog/no-code-lead-generation Published: 2025-11-05 # No-Code Lead Generation - How to Build a Lead Generation System Without Programming Lead acquisition is one of the most critical processes in any B2B company. But manually searching for contacts is incredibly time-consuming. Let me show you how to build a system that does it automatically. ## The Problem: Manual Lead Generation A typical manual lead acquisition process: 1. Searching for companies on Google (15-30 min per list) 2. Visiting company websites, looking for contacts (5-10 min per company) 3. Verifying email addresses (2-3 min per contact) 4. Entering data into CRM (1-2 min per record) **Result:** 2-3 hours of work = 10-15 verified leads ## The Solution: Automated Lead Generator A system built on: - **n8n** - workflow automation - **Snov.io** - email finder & verifier - **Apollo** - company and contact database - **The Company API** - company data - **Airtable** - central lead database ## How Does It Work? ### 1. Defining Search Criteria In Airtable, we create a "Campaigns" table where we define: - Industry (e.g., "e-commerce", "SaaS") - Company size (employees, revenue) - Location - Decision-maker titles (CEO, CTO, Marketing Director) ### 2. Automated Company Search ```text n8n workflow: 1. Trigger: New campaign in Airtable 2. Apollo Search: find companies matching criteria 3. The Company API: enrich company data 4. Filtering: remove duplicates and outdated data 5. Save to Airtable ``` ### 3. Finding Decision-Maker Contacts For each company, the system: - Searches for decision-maker profiles on LinkedIn (Apollo) - Finds email addresses (Snov.io) - Verifies email validity (Snov.io) - Adds contacts to the database ### 4. Data Enrichment The system automatically adds: - Company size - Revenue (if available) - Technologies used on the website - Social media activity - Recent company news ## System Architecture **Airtable Database:** - "Campaigns" table - campaign definitions - "Companies" table - discovered companies - "Contacts" table - decision-makers at companies - "Enrichment" table - additional data **n8n Workflows:** - Company Search (runs once daily) - Contact Finder (continuous, for new companies) - Email Verifier (verification every 7 days) - Data Enrichment (data enrichment) ## Efficiency **Manually (8h of work):** - 30-40 companies - 60-80 verified contacts - Cost: 8h x hourly rate **Automatically (24h):** - 500-1,000 companies - 1,500-3,000 verified contacts - Cost: ~$50 (API calls + tools) **ROI: 95% reduction in cost per lead** ## Operating Costs Monthly costs for a sample setup: - n8n (self-hosted): $0 - Snov.io (1,000 credits): $39/mo - Apollo (Basic): $49/mo - The Company API (1,000 calls): $29/mo - Airtable (Pro): $20/mo **Total: ~$140/mo** vs hundreds of hours of manual work ## Case Study: Marketing Agency **Before:** - 2 people full-time on prospecting - ~200 leads/month - Cost: 2 x $3,000 = $6,000/mo **After implementation:** - Automated system - ~2,500 leads/month - 1 person part-time on verification - Cost: $140 (tools) + $1,000 (partial salary) = $1,140/mo **Savings: $4,860/mo (81%)** ## Step-by-Step Implementation 1. **Week 1:** Tool configuration and integrations 2. **Week 2:** Building workflows in n8n 3. **Week 3:** Testing and optimization 4. **Week 4:** Training and launch Implementation time: 3-4 weeks ## Summary The Lead Generator is a system that: - Works 24/7 without breaks - Generates 10-15x more leads - Costs 80-90% less than manual work - Delivers higher quality data

Want to automate your lead generation?

I'll help you build an automated lead generation system that works 24/7, finds potential customers, and qualifies them before first contact. From concept through configuration to optimization.

Book a free consultation
## FAQ
### What tools make up a no-code automated lead generation system? The tech stack includes: n8n (workflow automation, self-hosted for $0), Snov.io (email finder and verification), Apollo (company and contact database), The Company API (company data), and Airtable (central lead database). All tools integrate without writing code through APIs and ready-made connectors.
### What savings does lead generation automation deliver compared to manual work? A 95% reduction in cost per lead and 10-15x more leads. Manually: 8h of work = 30-40 companies and 60-80 contacts. Automatically: 24h of system operation = 500-1,000 companies and 1,500-3,000 verified contacts for ~$50 in API costs. The system runs 24/7 without breaks.
### How does automated company and decision-maker contact search work? You define criteria in Airtable (industry, company size, location, job titles), n8n triggers the workflow: Apollo searches for companies, The Company API enriches the data, the system filters duplicates, and Snov.io finds and verifies decision-maker emails. Contacts land in the central database ready for outreach.
### How much does an automated lead generation system cost per month? About $140/month: n8n self-hosted $0, Snov.io (1,000 credits) $39, Apollo Basic $49, The Company API (1,000 calls) $29, Airtable Pro $20. For comparison: 2 people full-time on manual prospecting costs $6,000/month at a much smaller scale.
### How long does it take to implement an automated lead generation system? 3-4 weeks: tool configuration and integrations (week 1), building workflows in n8n (week 2), testing and optimization (week 3), training and production launch (week 4). After implementation, the system requires minimal maintenance - a part-time role for lead quality verification.
--- # AI-Powered Chatbots - From Concept to Deployment Source: https://pawel.lipowczan.pl/en/blog/ai-chatbots-from-concept-to-deployment Published: 2025-11-01 # AI-Powered Chatbots - From Concept to Deployment Next-generation chatbots powered by LLMs (Large Language Models) can hold natural, context-aware conversations. They're no longer limited to rigid scripts -- they understand intent and adapt to context. ## How Do They Differ from Traditional Chatbots? **Traditional chatbots:** - Rigid scenarios (decision trees) - Keyword matching - No understanding of context - Limited flexibility **AI chatbots:** - Natural language understanding - Context-aware responses - Conversation memory - Action execution (booking, search, etc.) ## Architecture of a Context-Based Chatbot ### 1. Frontend - User Interface **Web widget:** - Embedded on the website - Responsive design - Multimedia options (text, images, buttons) **Voicebot:** - Phone (VAPI) - Voice interface on the website - IVR integration ### 2. Backend - Conversation Logic (n8n) ```text n8n Workflow: 1. Webhook receive message 2. Load conversation context 3. Search knowledge base (RAG) 4. Call LLM (OpenAI/Claude) 5. Execute actions if needed 6. Store conversation history 7. Return response ``` ### 3. Knowledge Base - Source of Truth **RAG (Retrieval Augmented Generation):** Instead of training the model on your data, we use RAG: 1. Documents are split into chunks 2. Chunks are embedded (vectors) 3. Stored in a vector database (Qdrant) 4. On query: semantic search -> top N chunks -> context for LLM **Benefits of RAG:** - Up-to-date knowledge (update without retraining) - Lower costs - Better source attribution - Full data control ### 4. LLM - The Brain of the System **OpenAI GPT-4:** - Best response quality - Function calling (actions) - Cost: ~$0.01 per 1k tokens **Claude 3.5 Sonnet:** - Excellent at analysis - Large context (200k tokens) - Cost: ~$0.003 per 1k tokens ## Step-by-Step Implementation ### Step 1: Preparing the Knowledge Base Gather documents: - FAQ - Product documentation - Blog articles - Company policies Processing: ```python # Split into chunks (500-1000 tokens) # Embed via OpenAI ada-002 # Store in Qdrant ``` ### Step 2: Configuring the n8n Workflow **Main conversation flow:** 1. Webhook trigger (user message) 2. Vector search in Qdrant (top 3 relevant chunks) 3. Format prompt with context 4. Call OpenAI with function calling 5. If function -> execute & respond 6. Save to conversation history ### Step 3: Function Calling - Actions The chatbot can execute actions: ```json { "name": "book_meeting", "description": "Books a meeting with sales team", "parameters": { "date": "2025-11-20", "time": "14:00", "email": "user@example.com" } } ``` n8n detects the function call -> integrates with Calendly/Google Calendar -> confirmation ### Step 4: Testing and Optimization - Test different prompts - Analyze failed conversations - A/B test responses - Monitor accuracy ## Case Study: automation.house **Challenge:** The automation.house website had many offerings (Note Taker, Lead Generator, etc.). Users struggled to choose the right solution. **Solution:** A context-based chatbot that: - Asks questions about the client's needs - Understands the business context - Recommends appropriate solutions - Schedules consultations **Stack:** - n8n (hosting + workflow) - OpenAI GPT-4o (conversation) - Qdrant (product knowledge base) - Airtable (conversation tracking) **Results:** - 40% increase in engagement - 25% more consultations booked - 80% of users complete the conversation with a specific action ## Voicebots with VAPI VAPI is a platform for building voice AI: **Features:** - Real-time voice conversations - Telephony integration - Transfer to a human agent - Recording & transcription **Use cases:** - Automated helpline - Phone-based lead qualification - 24/7 customer support - Appointment booking ## Deployment Costs **Setup (one-time):** - Knowledge base preparation: 1-2 weeks - Workflow configuration: 1 week - Testing: 1 week - **Total: 3-4 weeks** **Monthly operating costs:** - n8n (self-hosted): $0-20 - OpenAI API (1,000 conversations): $30-50 - Qdrant Cloud: $25 - VAPI (voicebot): $99 - **Total: $150-200/mo** vs. 1 customer support employee: $2,500-3,500/mo ## Best Practices 1. **Clear conversation goal** - the bot must know what it's trying to achieve 2. **Graceful degradation** - transfer to a human when unsure 3. **Short responses** - don't write essays 4. **Personality** - give the bot a character that aligns with your brand 5. **Testing** - test with real users ## Summary Context-based chatbots are the future of customer experience: - 24/7 availability - Consistent quality - Scalability - Low operating cost

Want to deploy an AI chatbot in your company?

I can help you design, build, and deploy a chatbot tailored to your business needs. From use case analysis through knowledge base configuration to integration and optimization.

Book a free consultation
## FAQ
### What is RAG and why is it better than fine-tuning an AI model? RAG (Retrieval Augmented Generation) is a technique where the chatbot searches a knowledge base and provides the found information as context for the LLM. Its advantages over fine-tuning include: updating knowledge without costly retraining, lower costs, better control over response sources, and the ability to cite where information comes from.
### How much does it cost to deploy and maintain an AI chatbot for a business? Setup takes 3-4 weeks (knowledge base, workflow, testing). Monthly operating costs for 1,000 conversations: n8n self-hosted $0-20, OpenAI API $30-50, Qdrant Cloud $25, optionally VAPI for voicebot $99. Total $150-200/month vs $2,500-3,500 for a customer support employee.
### How do AI chatbots differ from traditional keyword-based chatbots? Traditional chatbots run on rigid scenarios (decision trees) and match keywords. AI chatbots understand natural language, remember conversation context, adapt to user intent, and execute actions (bookings, searches). The difference is the scale of flexibility -- AI handles queries its creator never anticipated.
### How do you prepare a knowledge base for a RAG-based chatbot? Gather documents (FAQ, product documentation, articles, company policies), split them into 500-1,000 token chunks, generate embeddings via OpenAI ada-002, and store them in a vector database (e.g., Qdrant). For each query, the chatbot retrieves the 3-5 most relevant fragments as context for its response.
### When should an AI chatbot hand the conversation over to a human? When it doesn't know the answer, the user is frustrated, the matter requires decisions beyond the bot's authority, or it involves sensitive topics (complaints, legal issues). Graceful degradation is a key best practice -- the bot informs the user it's transferring them to a consultant, rather than generating uncertain responses.
--- # Co to jest drugi mózg i po co Source: https://pawel.lipowczan.pl/llm-wiki/kurs/0-co-to-drugi-mozg Cel tej lekcji: zrozumieć ideę drugiego mózgu bez ani jednego technicznego słowa. Po lekcji wiesz, czym jest drugi mózg, po co go budować i że **nie trzeba umieć programować**. ## Znasz to uczucie Coś już kiedyś przeczytałeś albo rozwiązałeś. Miesiąc później szukasz tego od zera - w mailach, w głowie, w internecie. Robisz tę samą robotę drugi raz. ## Drugi mózg to lekarstwo na to Drugi mózg to zewnętrzna pamięć na rzeczy, które warto zatrzymać. Nie w głowie, tylko w plikach, do których zawsze wrócisz: rozwiązany problem, wnioski, sprawdzone podejście, ważna decyzja. ## Nowość: korzysta z niego też AI Z tej pamięci korzystasz nie tylko Ty. Korzysta z niej też asystent AI (program, który rozumie polecenia po ludzku i sam wykonuje zadania w Twoim imieniu). Czyta Twoje notatki i działa z Twoim kontekstem: Twoimi decyzjami, Twoim stylem, tym, co już kiedyś ustaliłeś. ## A skąd nazwa „LLM Wiki”? To po prostu sposób ułożenia tych notatek tak, żeby asystent AI szybko trafiał do właściwej rzeczy, zamiast czytać wszystko naraz. Nazwa brzmi technicznie, a idea jest prosta: uporządkowana pamięć, którą rozumie i człowiek, i maszyna. ## Po co Ci to Przestajesz płacić drugi raz - czasem i uwagą - za wiedzę, którą raz już zdobyłeś. Wiedza się kumuluje, zamiast wyparowywać. ## Dla kogo Dla każdego, kto dużo czyta, ustala i rozwiązuje - i nie chce tego gubić. Handlowiec, konsultant, właściciel firmy. **Nie trzeba umieć programować.** ## Czego NIE musisz umieć Nie musisz znać terminala (czarnego okna, w którym wpisuje się polecenia tekstem), kodu ani gita (narzędzia, które pilnuje historii zmian w plikach). Jak odpalić to wszystko bez nich - pokazuję w lekcji 3. --- # Trzy pojęcia zanim zaczniesz Source: https://pawel.lipowczan.pl/llm-wiki/kurs/0-trzy-pojecia Cel tej lekcji: rozbroić trzy słowa, które padną w kursie. Po lekcji wiesz, co znaczy agent AI, repozytorium i markdown (i przy okazji „komenda”). ## Agent AI (asystent AI) Program, który rozumie polecenia po ludzku, czyta Twoje pliki i **sam wykonuje zadania** - na przykład „zrób notatkę z tego tekstu”. Różnica od zwykłego czatu: agent ma dostęp do Twoich plików i działa, a nie tylko rozmawia. ## Repozytorium (repo) Brzmi groźnie, a to po prostu **folder z plikami** - Twoja baza. Czasem ma kopię w internecie (na GitHub, czyli w serwisie do przechowywania takich folderów), żeby nic nie zginęło. To Twoja kopia zapasowa. Tyle. Gdy w kursie pada „repo”, myśl „folder”. ## Markdown Zwykły tekst z bardzo prostym formatowaniem: `#` robi nagłówek, `-` robi listę, `**tekst**` pogrubia. Otwierasz go w dowolnym edytorze. Nic więcej - żadnego kodu. ## Bonus: komenda i „/slash” Komenda to skrót, który uruchamia gotowe zadanie. Wpisujesz `/ingest`, a asystent wie, co zrobić - trochę jak makro albo skrót klawiszowy. Kurs używa kilku takich komend; poznasz je po kolei. ## To wszystkie „trudne” słowa Reszta kursu to już praktyka na gotowym szablonie (przygotowanej z góry bazie, którą tylko kopiujesz do siebie). --- # Uruchom w swoim narzędziu Source: https://pawel.lipowczan.pl/llm-wiki/kurs/0-uruchom-w-swoim-narzedziu Cel tej lekcji: pokazać, gdzie i jak to odpalić - bez terminala (czarnego okna z poleceniami) i bez komend gita. Po lekcji wiesz, że działa w Twoim narzędziu, jak wołać komendy, jak zapisać zmiany i pobrać szablon bez komend gita oraz po co git i tak warto mieć pod spodem. ## Jeden warunek, reszta to szczegóły Wystarczy narzędzie, w którym **asystent AI ma dostęp do Twoich plików**. Asystent AI to program, który czyta Twoje pliki i sam wykonuje zadania. Jak ma dostęp - działasz. Nie musisz używać terminala. Nie musisz pisać komend gita. ## Gdzie to działa W każdym z tych narzędzi wpisujesz polecenia w oknie czatu asystenta: - **Claude Code** - narzędzie, na którym powstał kurs. - **Claude Desktop** (zakładka „Code”) - to samo, ale w aplikacji z okienkami, bez terminala. - **GitHub Copilot** (w edytorze VS Code) - bardzo popularny. - **Codex** - wielu użytkowników pracuje właśnie na nim. - **Cursor** i **Antigravity** - edytory z asystentem w bocznym panelu (jeden z użytkowników prowadzi na tym cały kurs). ## Jak wołasz komendy Komenda to skrót, który uruchamia gotowe zadanie. W Claude Code i Claude Desktop (Code) piszesz wprost `/onboard`, `/ingest`. W pozostałych narzędziach albo tak samo w czacie, albo prosisz zwykłym zdaniem: *„wykonaj instrukcje z pliku `.claude/commands/ingest.md`”*. Efekt ten sam - bo te komendy to po prostu pliki z instrukcjami, które asystent czyta. ## Zapis bez komend gita - trzy drogi Git to narzędzie, które zapisuje historię zmian w plikach. Nie musisz znać jego komend. Masz trzy drogi: 1. Powiedz asystentowi: *„zapisz zmiany”* - zrobi to za Ciebie. 2. Kliknij przyciski w panelu. Narzędzia oparte na edytorze VS Code (Copilot, Cursor, Antigravity) mają gotowy panel z guzikami. 3. W trybie chmury asystent sam przygotuje zmiany do zatwierdzenia jednym kliknięciem. ## Po co w ogóle git, skoro nie wpisujesz jego komend Nawet gdy nie znasz ani jednej komendy, git pracuje pod spodem i daje Ci trzy rzeczy, które bardzo się przydają - zwłaszcza gdy pliki zmienia za Ciebie asystent: - **Widzisz, co się zmieniło.** Git pokazuje diff (zestawienie „przed i po”): dokładnie te linie, które asystent dodał albo skasował. Zerkasz i wiesz, co zrobił, zamiast wierzyć na słowo. - **Cofasz zmiany.** Coś poszło nie tak? Wracasz do wcześniejszej wersji jednym ruchem. Nic nie ginie na stałe. - **Masz historię.** Każdy zapis to punkt w czasie - co i kiedy się zmieniło. Możesz przewinąć bazę do dowolnego wcześniejszego momentu. To jak automatyczna kopia zapasowa z maszyną czasu. Wszystko obsłużysz przyciskami w panelu albo prosząc asystenta zdaniem - komendy gita są opcjonalne, korzyści nie. ## Nie chcesz w ogóle gita na start? Wejdź na stronę szablonu na GitHub: [github.com/plipowczan/second-brain-template](https://github.com/plipowczan/second-brain-template). Kliknij zielony przycisk **„Code” → „Download ZIP”**, rozpakuj folder i otwórz go w swoim narzędziu. Tyle. ## Uczciwa granica Jedyne, co NIE zadziała, to zwykła aplikacja-czat **bez dostępu do Twoich plików**. Wtedy albo wybierz jedno z narzędzi wyżej, albo podłącz do niej folder - ale to już bardziej techniczne i na start niepotrzebne. Gotowe? W następnej lekcji zakładasz katalog z szablonu i ruszamy z praktyką. --- # Załóż katalog z szablonu Source: https://pawel.lipowczan.pl/llm-wiki/kurs/1-zaloz-katalog Cel tej lekcji: przejść od „Use this template” do gotowej, uzbrojonej struktury bazy. Po lekcji rozumiesz koncept LLM Wiki, masz własne repo z szablonu i wiesz, **dlaczego** ta baza działa bez embeddings (liczbowych reprezentacji tekstu do wyszukiwania) i bez RAG (techniki, w której model przy każdym pytaniu przeszukuje surowe dokumenty) - architektura trzech warstw, trzy indeksy i zasada progressive disclosure (czytania od ogółu do szczegółu). ## Po co to LLM Wiki (koncept Karpathy'ego) odwraca RAG: zamiast za każdym razem przeszukiwać surowe dokumenty, agent **przyrostowo buduje i utrzymuje** żywą bazę markdown - z linkami, indeksami, syntezami. Wiedza **kumuluje się** i jest czytelna **i dla agenta, i dla człowieka** (otwierasz plik, czytasz). Standard **OKF** (Open Knowledge Format, Google) formalizuje ten wzorzec → bazy są przenośne. ## Weź szablon Na GitHubie „Use this template” (albo `git clone`) → własne repo. Otwórz folder w Claude Code. Opcjonalnie `npm install` (Prettier - formatowanie) i `pip install -r requirements.txt` (skille pythonowe) - **do pierwszego pytania niepotrzebne**. Publikację przez Quartz doklejasz później (lekcja 5). Repo szablonu: [github.com/plipowczan/second-brain-template](https://github.com/plipowczan/second-brain-template). ## Krok po kroku: z szablonu do VS Code Cały proces masz na wideo u góry. Poniżej to samo w krokach - ścieżka przez przeglądarkę + GitHub Desktop (najprościej dla każdego). **1. Wejdź na repo szablonu.** `github.com/plipowczan/second-brain-template`. Zielony przycisk **„Use this template”** w prawym górnym rogu. ![Strona szablonu second-brain-template na GitHub z przyciskiem Use this template](/images/kurs/1-zaloz-katalog-01.webp) **2. „Use this template” → „Create a new repository”.** Właściciel = Ty, **Repository name** = np. `brain-test`, widoczność Public lub Private (Twój wybór). Kliknij „Create repository”. ![Formularz tworzenia nowego repozytorium z szablonu - pole Repository name](/images/kurs/1-zaloz-katalog-02.webp) Po chwili masz **własne repo** z całą zawartością szablonu (tu: `brain-test`): ![Nowe repozytorium brain-test utworzone z szablonu](/images/kurs/1-zaloz-katalog-03.webp) **3. Sklonuj repo lokalnie.** Zielony **„Code” → „Open with GitHub Desktop”** (albo skopiuj URL HTTPS i użyj `git clone`). ![Menu Code z opcją Open with GitHub Desktop](/images/kurs/1-zaloz-katalog-04.webp) **4. W GitHub Desktop** potwierdź URL repo i **Local path** (np. `C:\Projects\brain-test`), potem **Clone**. ![Okno Clone a repository w GitHub Desktop - URL i ścieżka lokalna](/images/kurs/1-zaloz-katalog-05.webp) **5. Otwórz w edytorze.** Po sklonowaniu GitHub Desktop proponuje **„Open in Visual Studio Code”** (Ctrl+Shift+A). ![GitHub Desktop po klonie - przycisk Open in Visual Studio Code](/images/kurs/1-zaloz-katalog-06.webp) **6. Gotowe - VS Code otwarty na `brain-test`.** W Explorerze widzisz całą strukturę szablonu. Co jest czym - niżej. ![VS Code otwarty na repozytorium brain-test ze strukturą plików](/images/kurs/1-zaloz-katalog-07.webp) ## 🤖 Gotowy prompt - niech agent założy repo Ścieżka przez przeglądarkę („Use this template” wyżej) jest najprostsza. Jeśli wolisz, żeby **Claude Code sam** utworzył Twoje repo i sklonował je lokalnie - otwórz Claude Code w folderze na projekty i wklej (podmień ``): ```text Utwórz moje własne repo z szablonu second-brain i sklonuj je tutaj. Szablon: github.com/plipowczan/second-brain-template. Jeśli mam zalogowane gh CLI, użyj: gh repo create --template plipowczan/second-brain-template --private --clone W przeciwnym razie zrób git clone szablonu do folderu i wyjaśnij, jak później podmienić origin na moje własne repo. Po sklonowaniu wejdź do folderu, potwierdź strukturę (CLAUDE.template.md, content/_raw/inbox/, content/_indexes/, .claude/commands/) i powiedz, czy mogę odpalić /onboard. Poza klonowaniem niczego nie zmieniaj. ``` Agent założy repo, sklonuje, zweryfikuje strukturę i da zielone światło na lekcję 2. Tak to wygląda w praktyce - prompt odpalony w terminalu Claude Code (wklejenie, podmiana ``, uruchomienie, efekt w VS Code): ## Co jest w repo (po utworzeniu z szablonu) Świeży klon wygląda tak (to samo drzewo widać w Explorerze VS Code wyżej): ```text brain-test/ .claude/ commands + skills + hooks - mózg agenta .githooks/ pre-commit - higiena przed commitem content/ Twój vault (rozwinięcie niżej) tests/ run_tests.py - testy skryptów skilli CLAUDE.template.md · AGENTS.template.md schema agenta ({{placeholdery}}) content/WRITING_STYLE.template.md szablon Twojego głosu schema.yml kontrakt frontmatteru per typ noty README.md · package.json · requirements.txt .editorconfig · .gitattributes · .gitignore · .npmrc · .prettierrc ``` **Foldery:** - **`.claude/`** - mózg agenta. `commands/` to slash-komendy (`/onboard`, `/ingest`, `/qa`, `/lint`, `/reindex`, `/gaps`, `/curate`, `/compile`, `/enhance`, `/output`, `/refactor`), `skills/` to ich implementacje (m.in. pythonowe: reindex, lint, ingest, gaps + skille research i `excalidraw-diagram`). `hooks/load_vault_map.py` auto-ładuje mapę bazy na starcie każdej sesji, a `settings.json` to konfiguracja Claude Code. - **`.githooks/`** - hook `pre-commit`: kontrola jakości (format/lint), zanim cokolwiek trafi do commita. - **`content/`** - Twój vault (tak nazywamy folder z całą Twoją wiedzą): noty, indeksy, szablony i wzorce (rozbicie niżej). - **`tests/`** - `run_tests.py`, testy skryptów pythonowych skilli. Nie ruszasz na co dzień - są, żeby mechanika była pewna. **Pliki schemy (znikają po `/onboard`):** - **`CLAUDE.template.md` / `AGENTS.template.md`** - kontrakt agenta z placeholderami, czyli polami do podmiany (`{{KB_NAME}}`, `{{KB_OWNER}}`, `{{TOPIC_TABLE}}`, ...). `/onboard` wypełnia je i zapisuje jako `CLAUDE.md` / `AGENTS.md`. Ich obecność to znak, że baza jest jeszcze nieskonfigurowana. - **`content/WRITING_STYLE.template.md`** - szablon Twojego „głosu” (osoba, formalność, emoji w nagłówkach). `/onboard` wypełnia go i zapisuje jako `content/WRITING_STYLE.md`. **Pozostałe pliki:** - **`schema.yml`** - kontrakt frontmatteru (bloku metadanych na początku pliku) per typ noty: jakie pola są wymagane (np. `knowledge-note` musi mieć `title, date, type, tags, summary`). `/lint` to egzekwuje. Skasujesz plik → luźniejszy tryb, bez wymuszania pól. - **`README.md`** - opis szablonu i quickstart. - **`package.json`** - jedyny skrypt to `npm run format` (Prettier). Potok publikacji (Quartz) świadomie **nie** wchodzi w zestaw - doklejasz go później (lekcja 5). - **`requirements.txt`** - zależności Pythona skryptów (`PyYAML`; `yt-dlp` opcjonalnie, pod `/ingest` z YouTube). - **`.editorconfig` · `.gitattributes` · `.gitignore` · `.npmrc` · `.prettierrc`** - standardowa higiena repo (styl kodu, końcówki linii, ignorowane pliki, format). **W środku `content/`:** - **`_raw/inbox/`** - miejsce na surowe źródła (jest gotowy `sample-source.md` na rozgrzewkę). Po `/ingest` źródła lądują w `_raw/processed/`. - **`_indexes/`** - trzy auto-utrzymywane indeksy nawigacyjne: `vault-map.md`, `catalog.md`, `graph.md` (serce progressive disclosure - patrz niżej). - **`_outputs/`** - `answers/` (zapisane odpowiedzi `/qa`) i `reports/` (raporty `/lint`). - **`_graveyard/`** - wycofane noty: odwracalne, wykluczone z indeksów (`/curate` je tam odkłada). - **`templates/`** - szablony not per typ: `basic_notes.md`, `book.md`, `knowledge_note_info.md`, `knowledge_note_how_to.md`, `tool.md`. - **`REFERENCE/`** - wzorce na start: `Example Note.md` i `Wikilinks Explained.md` (jak wygląda dobra nota i jak działają `[[linki]]`). To „**pusty, ale uzbrojony**” brain - cała mechanika gotowa, brak tylko Twojej wiedzy. ## Architektura - 3 warstwy | Warstwa | Gdzie | Kto włada | | ----------- | ------------------------------- | ------------------------------------ | | Raw sources | `_raw/inbox` → `_raw/processed` | Ty wrzucasz, agent tylko czyta | | Wiki | `content//` | Agent tworzy/utrzymuje noty | | Schema | `CLAUDE.md` (z onboardingu) | Ty + agent współ-ewoluujecie | ## Dlaczego to działa bez RAG: progressive disclosure **Progressive disclosure** to jedna zasada: **najpierw pokaż, co istnieje i ile kosztuje pobranie - niech agent sam zdecyduje, co wczytać.** Tak jak człowiek skanuje spis treści przed rozdziałem albo nazwy plików, zanim któryś otworzy. Porównaj dwa podejścia do tego samego pytania: - **Klasyczny RAG - „wrzuć wszystko”.** Do kontekstu ląduje np. 35 000 tokenów notatek i historii, z czego realnie trafne jest może 2 000. Efektywność ~6%. Reszta zabiera uwagę modelu i miejsce na właściwe zadanie. - **Index-first - progressive disclosure.** Agent czyta najpierw **indeks** (~kilkaset tokenów): co istnieje, jakiego typu, z jakimi tagami. Na tej podstawie pobiera **tylko** 2-3 trafne noty (~kilkaset tokenów). Trafność bliska 100%, a okno kontekstu zostaje wolne na myślenie. To dlatego baza działa **bez embeddings i bez RAG do ~500 źródeł**: nie trzeba liczyć wektorów, gdy dobry indeks pozwala agentowi zawęzić wyszukiwanie samą lekturą. ## 3 indeksy - trzy poziomy powiększenia Trzy indeksy w `content/_indexes/` to ten sam vault widziany z trzech odległości. Agent zaczyna z lotu ptaka i **przybliża** dopiero tam, gdzie trzeba. ### vault-map.md - mapa z lotu ptaka (L0) Co gdzie leży: tabela folderów (ile not, jakie typy, dominujące tagi), chmura tagów i lista ostatnich zmian. Jedna „strona”, z której agent wybiera **folder**, nie notę. ```text --- updated: 2026-06-30T12:00:00Z total_notes: 330 --- # Vault Map ## Folders | folder | notes | types | top-tags | |--------------------|------:|--------------------|-------------------------------------| | AI/KNOWLEDGE/INFO | 28 | knowledge-note(27) | ai, agents, context-engineering | | AI/TOOLS | 69 | tool(69) | tool, claude-code, agents | | CODE/TOOLS | 40 | tool(40) | frontend, framework, infrastructure | ## Tag Cloud agents:29 ai:119 claude-code:38 context-engineering:11 rag:10 skills:16 ... ## Recent Changes - 2026-06-30 AI/TOOLS/Ponytail (ingest - vibe-coding skill: YAGNI-first, −54% LOC) - 2026-06-25 CODE/TOOLS/Git (enhance - uzupełniony stub + linki) ``` Czytając to, agent od razu wie: „temat o agentach i tokenach? → folder `AI/KNOWLEDGE/INFO`”. Zawęża, zanim cokolwiek otworzy. ### catalog.md - jedna linia na notę (L1) Katalog wszystkich not, pogrupowany folderami. Każda nota to **jedna linia**: tytuł, typ, data, tagi, jednozdaniowe streszczenie i lista wychodzących linków. ```text ## AI/KNOWLEDGE/INFO - **Context Engineering** | knowledge-note | 2026-04-09 | [ai, context-engineering] | Zarządzanie tym, co trafia do okna kontekstu agenta - indeksy, kompresja, progressive disclosure. | → Progressive Disclosure, Agent Skills, Claude Code - **Progressive Disclosure** | knowledge-note | 2026-04-30 | [ai, memory] | Pokaż najpierw co istnieje i ile kosztuje pobranie; agent sam decyduje, co wczytać. | → Context Engineering, Harness Engineering ``` Format linii: `**Tytuł** | typ | data | [tagi] | streszczenie | → linki`. Po streszczeniach agent wybiera **konkretne noty** do otwarcia - wciąż nie czytając ich treści. ### graph.md - graf połączeń (L2) Kto z kim linkuje. Dzięki temu agent po otwarciu jednej noty wie, które sąsiednie **dociągnąć**, żeby spiąć temat. ```text ## Outgoing AI/KNOWLEDGE/INFO/Context Engineering -> AI/KNOWLEDGE/INFO/Progressive Disclosure, AI/TOOLS/Agent Skills, AI/TOOLS/Claude Code AI/KNOWLEDGE/INFO/Progressive Disclosure -> AI/KNOWLEDGE/INFO/Context Engineering, AI/KNOWLEDGE/INFO/Harness Engineering ``` ## Jak agent tego używa - przykład `/qa` Zapytanie: `/qa "jak ograniczyć zużycie tokenów w agencie?"` 1. **vault-map (L0):** agent czyta mapę (~jedna strona). Widzi, że temat pasuje do `AI/KNOWLEDGE/INFO` (tagi `context-engineering`, `token-optimization`). Zawęża do tego folderu. 2. **catalog (L1):** czyta tylko linie tego folderu. Po streszczeniach wybiera 2 noty: „Token Optimization for Claude Code” i „Progressive Disclosure”. 3. **Pełne noty (Layer 2):** otwiera **tylko te 2** - nie 330. 4. **graph (L2):** widzi, że obie linkują do „Context Engineering”, więc dociąga ją, bo spina temat. Efekt: agent przeczytał ~3 noty zamiast całej bazy. Kilkaset tokenów indeksu + kilka trafnych not - zamiast wrzucania wszystkiego. To jest progressive disclosure w akcji. ## Wzorzec - jak wygląda dobra nota Zajrzyj do `REFERENCE/Example Note.md` i `Wikilinks Explained.md` - jak wygląda dobra nota i jak działają linki. Dobra nota ma frontmatter z `type`, jednozdaniowe `summary` (to trafia do katalogu) i kilka celnych `[[wikilinków]]` (to zasila graf). ## Pułapka Nie pisz wiki ręcznie - to robota agenta; Ty dostarczasz źródła i dobre pytania. Nie edytuj też indeksów z palca - buduje je `/ingest` i `/reindex` (o tym w lekcjach 3 i 4). ## FAQ ### Czym LLM Wiki różni się od RAG i baz wektorowych? RAG za każdym pytaniem przeszukuje surowe dokumenty i wrzuca do kontekstu masę tekstu, z którego trafna jest garstka. LLM Wiki odwraca to: agent czyta najpierw lekki indeks (co istnieje, jakiego typu, z jakimi tagami) i dociąga tylko 2-3 trafne noty. Nie liczy wektorów ani embeddingów - do ~500 źródeł dobry indeks wystarcza, żeby zawęzić wyszukiwanie samą lekturą. ### Agent ma grep - po co mu jeszcze indeks? Grep (wyszukiwanie w plikach po dokładnym słowie) wystarcza, gdy baza jest mała, a Ty znasz szukane słowo - i agent wciąż go używa. Problem zaczyna się ze skalą: przy setkach not częste słowo zwraca dziesiątki trafień, agent wczytuje je wszystkie i okno kontekstu puchnie. Do tego grep nie zna synonimów - szukając „składki zdrowotnej”, nie znajdzie „ubezpieczenia zdrowotnego”; opisy i linki w indeksie niosą znaczenie, nie tylko ciąg znaków. Indeks nie zastępuje grepa: najpierw zawęża zakres do 2-3 właściwych not, a grep szuka już tylko wewnątrz nich. ### Czy do startu muszę odpalać `npm install` i `pip install`? Nie, do pierwszego pytania nie są potrzebne. `npm install` instaluje tylko narzędzia dev (Prettier - formatowanie), a `pip install -r requirements.txt` - zależności skryptów pythonowych (`/reindex`, `/lint`, `/ingest`, `render.py`). Publikację przez Quartz doklejasz osobno w lekcji 5. Sama struktura i noty to zwykły markdown - możesz zacząć od razu. ### Repo z szablonu - publiczne czy prywatne? Twój wybór; baza to zwykłe repo git i działa tak samo w obu trybach. Prywatne, jeśli to Twoja osobista wiedza (domyślnie bezpieczniej). Publiczne - jeśli od razu chcesz ją publikować lub dzielić się nią jako paczką OKF. Widoczność zmienisz w każdej chwili w ustawieniach repo. ### Czy potrzebuję płatnych narzędzi albo zewnętrznej bazy danych? Szablon jest darmowy, a cała baza to pliki markdown w Twoim repo - żadnej zewnętrznej bazy danych, wektorowej usługi ani API do przechowywania wiedzy. Potrzebujesz Claude Code do pracy z agentem oraz (opcjonalnie) Pythona do skryptowych skilli. Nic nie wychodzi na zewnątrz, dopóki sam nie opublikujesz. ### Co znaczy, że baza jest „pusta, ale uzbrojona”? Cała mechanika jest gotowa od pierwszej minuty: slash-komendy, skille, trzy indeksy nawigacyjne, hook ładujący mapę vaultu, kontrakt frontmatteru (`schema.yml`) i szablony not. Brakuje wyłącznie Twojej wiedzy. Nie budujesz narzędzi - od razu ich używasz, a baza rośnie z każdym źródłem. ### Co, gdy baza urośnie powyżej ~500 źródeł? Progressive disclosure skaluje się dalej - indeks i graf rosną wolniej niż treść, więc agent nadal zawęża zanim otworzy notę. Przy bardzo dużych bazach można podzielić wiedzę na tematyczne podbazy (multi-brain, lekcja 5) albo dołożyć warstwę wyszukiwania. Granica ~500 to moment, w którym warto o tym pomyśleć, a nie twardy limit. --- # Onboarding Source: https://pawel.lipowczan.pl/llm-wiki/kurs/2-onboarding Cel tej lekcji: przejść **cały onboarding krok po kroku** (onboarding = pierwsza konfiguracja bazy) na konkretnym przykładzie. Po lekcji masz wygenerowaną schema (`CLAUDE.md` - plik zasad Twojej bazy), foldery tematów i puste, ale gotowe indeksy - i wiesz dokładnie, co dzieje się pod spodem. Prowadzę Cię przez **realny przebieg** (screeny z prawdziwej sesji): **Anna** stawia bazę „Baza Anny” po polsku, na tematach `AI`, `BUSINESS`, `HEALTH`. Twoje odpowiedzi będą inne - proces jest ten sam. ## Warunek startu Onboarding odpalasz **raz**, na świeżym klonie. Kreator rozpoznaje świeży klon po obecności pliku `CLAUDE.template.md` w repo. Otwórz folder w Claude Code i wpisz: ```text /onboard ``` ![Sesja Claude Code otwarta na repo bazy (brain-test) z wpisanym /onboard](/images/kurs/2-onboarding-01.webp) Jeśli `CLAUDE.template.md` już nie istnieje (baza była konfigurowana), kreator nie nadpisze niczego - zaproponuje **ponowną konfigurację** (patrz „Powtarzalność” niżej). ## Faza 1 - Wywiad (jedno pytanie na raz) Kreator prowadzi **sześć pytań, pojedynczo** - wybierasz z gotowych opcji albo wpisujesz swoje. Po każdej odpowiedzi powtarza wybór i zapowiada następne pytanie, więc widzisz cały ślad decyzji. Poniżej realny przebieg „Bazy Anny”. **Q1 - Nazwa bazy i właściciel.** Wolny tekst; kreator podpowiada warianty (np. z Twojej tożsamości git). Anna wpisuje `Baza Anny · Anna Kowalska`. ![Pytanie 1 - tożsamość bazy: nazwa bazy i właściciel, z podpowiedziami](/images/kurs/2-onboarding-02.webp) **Q2 - Główny język.** Język, w którym agent pisze noty i streszczenia. Anna: **Polish**. ![Pytanie 2 - główny język: Polish / English / Mixed](/images/kurs/2-onboarding-03.webp) **Q3 - Tematy / domeny.** Foldery najwyższego poziomu pod `content/` - szerokie działy, nie szczegóły. Kreator proponuje gotowe zestawy; Anna wybiera **AI, BUSINESS, HEALTH**. ![Pytanie 3 - tematy: gotowe zestawy startowe albo własna lista](/images/kurs/2-onboarding-04.webp) **Q4 - Typy not.** Domyślnie 5: `basic-note`, `knowledge-note`, `tool`, `book-note`, `answer-note`. Anna zostawia **wszystkie 5**. ![Pytanie 4 - typy not: wszystkie 5 domyślnych albo własny zestaw](/images/kurs/2-onboarding-05.webp) **Q5 - Głos.** Osoba + formalność. Anna: **first-person, direct & practical** („Notuję”, ton casual-professional). ![Pytanie 5 - głos: osoba i formalność](/images/kurs/2-onboarding-06.webp) To samo pytanie ma drugą zakładkę - **emoji w nagłówkach** (`## 💡 Insight`). Anna zostawia **Yes**. ![Pytanie 5, druga zakładka - emoji w nagłówkach: tak/nie](/images/kurs/2-onboarding-07.webp) Głos + emoji zatwierdzasz razem - kreator pokazuje **przegląd** przed wysłaniem: ![Przegląd odpowiedzi o głosie i emoji przed zatwierdzeniem](/images/kurs/2-onboarding-08.webp) **Q6 - Główna gałąź.** Domyślnie `main` (pasuje do repo). Anna: **main**. ![Pytanie 6 - główna gałąź git: main / master](/images/kurs/2-onboarding-09.webp) ## Faza 2 - Generowanie (deterministyczne) Po ostatnim pytaniu kreator działa sam. Najpierw zapisuje odpowiedzi do `.kb-onboard.json` w korzeniu repo (to stąd bierze się jej powtarzalność): ```json { "KB_NAME": "Baza Anny", "KB_OWNER": "Anna Kowalska", "PRIMARY_LANGUAGE": "Polish", "MAIN_BRANCH": "main", "VOICE_PERSON": "first-person", "VOICE_FORMALITY": "direct and practical", "NOTE_TYPES": "basic-note | knowledge-note | tool | book-note | answer-note", "emoji_headings": true, "TOPIC_TABLE": "| `/content/AI/` | AI, agents, and tooling notes | Yes |\n| `/content/BUSINESS/` | Sales, marketing, operations notes | Yes |\n| `/content/HEALTH/` | Health, fitness, and wellbeing notes | Yes |" } ``` Potem, krok po kroku (widać to na screenie): 1. **Wypełnia trzy pliki schematu** z szablonów (skrypt `render.py`): `CLAUDE.md`, `AGENTS.md`, `content/WRITING_STYLE.md`. 2. **Kasuje szablony** (`*.template.md`) - dopiero gdy wszystkie trzy pliki wygenerują się poprawnie. 3. **Tworzy foldery tematów** pod `content/`: `AI/`, `BUSINESS/`, `HEALTH/`. 4. **Pyta o przykłady**: `content/REFERENCE/` (2 noty) + `content/_raw/inbox/sample-source.md` - usunąć czy zostawić jako samouczek. Anna wybiera **„Keep as tutorial”**. 5. **Przycina szablony not** do wybranych typów (tu wszystkie 5 zostają - nic nie ubywa). 6. **Przebudowuje indeksy** (`build_indexes.py`). ![Faza generowania: wypełnianie plików schemy, kasowanie szablonów, tworzenie folderów i pytanie o przykładowe treści](/images/kurs/2-onboarding-10.webp) ## Faza 3 - Gotowe (podsumowanie + przekazanie) Kreator drukuje podsumowanie: co powstało i co dalej. ![Podsumowanie onboardingu: wygenerowane pliki, foldery tematów, indeksy, sprawdzenie zależności i „co teraz”](/images/kurs/2-onboarding-11.webp) Dla „Bazy Anny” wyszło: - `CLAUDE.md`, `AGENTS.md`, `content/WRITING_STYLE.md` **wygenerowane**, szablony **skasowane**, - foldery tematów: `content/AI/`, `content/BUSINESS/`, `content/HEALTH/`, - typy not: wszystkie 5; głos first-person direct/practical; emoji w nagłówkach: on, - `REFERENCE/` + `sample-source` **zostawione jako samouczek**, - **indeksy przebudowane** - 2 noty, 4 krawędzie (to te 2 przykładowe noty), - **zależności**: Python 3.13 ✅, yt-dlp ✅, ffmpeg ✅ (Python wymagany do `reindex`/`lint`/`render`; yt-dlp+ffmpeg opcjonalne, pod `/ingest` z YouTube). Repo po fazie: ```text content/ AI/ .gitkeep BUSINESS/ .gitkeep HEALTH/ .gitkeep REFERENCE/ Example Note.md · Wikilinks Explained.md (zostawione jako samouczek) _indexes/ vault-map.md · catalog.md · graph.md (przebudowane: 2 noty, 4 krawędzie) templates/ (5 typów not) CLAUDE.md ← wygenerowany (schema Twojej bazy) AGENTS.md ← wygenerowany content/WRITING_STYLE.md ← wygenerowany (Twój głos) ``` Szablony `*.template.md` zniknęły - to znak, że baza jest zainicjalizowana. Kreator **nie robi żadnego commita** („Not committed - your call”). **Co teraz** (kreator podpowiada): - wrzuć plik do `content/_raw/inbox/` i odpal `/ingest` (lekcja 3), - zadaj pytanie przez `/qa` (lekcja 4), - odpal `/lint` na przegląd stanu bazy (lekcja 4). ## Zapisz stan (git) Baza to zwykłe repo git - pierwszy commit robisz jak zawsze. Możesz to nawet zlecić agentowi („commit this”): sam doda pliki i zrobi commit, **pomijając `.kb-onboard.json`** (jest pomijany przez gita - trzyma Twoje odpowiedzi lokalnie). ![Agent robi pierwszy commit świeżo skonfigurowanej bazy - czysty stan repo](/images/kurs/2-onboarding-12.webp) ## Powtarzalność Odpowiedzi siedzą w `.kb-onboard.json`. Gdy odpalisz `/onboard` ponownie na już skonfigurowanej bazie (brak `CLAUDE.template.md`), kreator **nie nadpisze niczego po cichu** - zaproponuje **ponowną konfigurację**: wczyta `.kb-onboard.json`, pozwoli poprawić odpowiedzi i przegeneruje pliki, ostrzegając przed nadpisaniem istniejących. ## 🤖 Szybka ścieżka - onboarding promptem `/onboard` prowadzi wywiad pytanie-po-pytaniu (jak wyżej). Jeśli wolisz, żeby Claude Code wykonał wszystko za jednym zamachem, wklej prompt z gotowymi odpowiedziami (podmień `<...>` na swoje): ```text Przeprowadź onboarding tej bazy (skill onboard). Moje odpowiedzi: - nazwa bazy / właściciel: / - język: polski - tematy: , , - typy not: domyślne - głos: pierwsza osoba, praktyczny, emoji w nagłówkach: tak - gałąź główna: main Wykonaj wszystkie fazy: wygeneruj CLAUDE.md / AGENTS.md / WRITING_STYLE.md, utwórz foldery tematów, skasuj szablony *.template.md, przytnij szablony not do wybranych typów i przebuduj indeksy. Na koniec pokaż drzewo repo i podsumowanie. ``` Efekt ten sam co wywiad - bez klikania. Odpowiedzi lądują w `.kb-onboard.json`, więc ponowna konfiguracja dalej działa. ## Dzień zerowy To Twój „**dzień zerowy**” - jedyny moment konfiguracji. Potem już tylko używasz: `/ingest`, `/qa`, `/lint`. Przechodzimy do karmienia bazy w lekcji 3. ## FAQ ### Ile razy odpalam `/onboard`? Raz, na świeżym klonie. Kreator rozpoznaje świeży klon po obecności pliku `CLAUDE.template.md`. Jeśli odpalisz go ponownie na skonfigurowanej bazie (szablonu już nie ma), nie nadpisze niczego po cichu - zaproponuje **ponowną konfigurację**. ### Co, jeśli wybiorę zły język albo tematy? Poprawiasz przez ponowny `/onboard` w trybie ponownej konfiguracji: kreator wczyta Twoje odpowiedzi z `.kb-onboard.json`, pozwoli je zmienić i przegeneruje pliki, ostrzegając przed nadpisaniem. Foldery tematów to zwykłe katalogi pod `content/` - nowe możesz dodać ręcznie w dowolnym momencie, a `/reindex` odświeży indeksy. ### Gdzie zapisują się moje odpowiedzi z wywiadu? W pliku `.kb-onboard.json` w korzeniu repo. Jest **pomijany przez gita** - trzyma Twoje wybory lokalnie i nie trafia do commita. To z niego bierze się powtarzalność: ponowna konfiguracja czyta właśnie ten plik. ### Dlaczego szablony `*.template.md` znikają po onboardingu? Bo kreator wypełnia je Twoimi wartościami i zapisuje jako finalne pliki: `CLAUDE.template.md` → `CLAUDE.md`, `AGENTS.template.md` → `AGENTS.md`, `content/WRITING_STYLE.template.md` → `content/WRITING_STYLE.md`. Szablony kasuje **dopiero gdy wszystkie trzy pliki wygenerują się poprawnie**. Ich brak to znak, że baza jest zainicjalizowana. ### Czy `/onboard` robi commit do gita? Nie - kończy komunikatem „Not committed - your call”. Pierwszy commit robisz sam (albo zlecasz agentowi: „commit this”). Agent doda pliki i zrobi commit, **pomijając `.kb-onboard.json`** (jest pomijany przez gita). ### Czy do onboardingu potrzebuję Pythona? Tak - kreator używa skryptów pythonowych (`render.py` do wygenerowania plików schemy, `build_indexes.py` do indeksów), a Python jest też wymagany później przez `/reindex` i `/lint`. `yt-dlp` i `ffmpeg` są opcjonalne - przydają się dopiero przy `/ingest` z YouTube (lekcja 3). --- # Pierwszy ingest Source: https://pawel.lipowczan.pl/llm-wiki/kurs/3-pierwszy-ingest Cel tej lekcji: zamienić surowe źródło w noty i indeksy. Po lekcji umiesz dokarmiać bazę nowymi elementami. ## Wrzuć źródło i odpal `/ingest` Wrzuć dowolne źródło do `content/_raw/inbox/` (jest gotowy `sample-source.md` na rozgrzewkę) i odpal **`/ingest`**. Tyle - komenda bierze wszystko, co leży w inboxie. `/ingest` przyjmuje też linki YouTube jako argument (`/ingest https://youtu.be/...`) - pobiera transkrypt (tekstowy zapis nagrania) i traktuje go jak zwykłe źródło. W przykładzie: surowe źródło to `sample-source.md` - luźny tekst o metodzie **Zettelkasten**. Ingest rusza, czyta plik, widzi pojedyncze źródło (brak klastra, czyli innych plików o tym samym temacie) i ustala temat oraz typ noty. ![Surowe źródło (sample-source.md o Zettelkasten) i ingest w akcji: czyta plik, brak klastra, ustala temat i typ](/images/kurs/3-pierwszy-ingest-02.webp) ## Co się dzieje pod spodem (4 fazy) `/ingest` nie „wkleja pliku do wiki”. Prowadzi źródło przez cztery fazy - trzy autonomiczne, jedna z jednym pytaniem do Ciebie. ### Faza 0 - Pobranie z YouTube (tylko dla URL) Gdy podasz link YouTube, agent najpierw ściąga transkrypt (`yt-dlp`; przy braku napisów - awaryjnie Whisper przez `ffmpeg`) i archiwizuje go w `content/_raw/processed/`. Dla zwykłych plików ta faza jest pomijana. Jeśli to samo wideo już wczytałeś (ten sam `video_id`), agent nie pobiera drugi raz - dołoży do istniejącej noty. ### Faza 1 - Wstępne rozpoznanie Zanim cokolwiek utworzy, agent: 1. czyta `vault-map.md` - żeby wiedzieć, co w bazie już jest (nie duplikować), 2. listuje inbox (+ ewentualne źródła z Fazy 0), 3. **wykrywa klastry** - pliki o wspólnym temacie (≥2 wspólne, wyróżniające słowa w tytułach/treści). Jeśli pliki układają się w grupę, dostajesz **jedno pytanie**: potraktować je jako osobne noty, czy jako notę nadrzędną + noty-dzieci z linkami. To jedyny moment, w którym ingest Cię pyta - reszta idzie sama. Po co: żeby powiązane źródła (np. 5 plików o jednym produkcie) nie rozsypały się na oderwane noty bez wspólnego punktu. ### Faza 2 - Wykonanie (autonomiczne) Dla każdego źródła agent po kolei: 1. **Klasyfikuje** - dobiera folder tematyczny i typ noty (`knowledge-note`, `tool`, `book-note`...) wg reguł z `CLAUDE.md`. Gdy nie jest pewny - daje najlepszy strzał + tag `#todo/classification`, żebyś to przejrzał. 2. **Sprawdza pokrywanie** z `catalog.md`: jest podobna nota → **scalenie** (dokłada, nie nadpisuje Twojej treści); brak → **nowa nota** z szablonu z `content/templates/`. 3. **Wypełnia frontmatter** (blok metadanych na początku pliku): `title`, `date`, `tags`, **`type`** (minimum OKF, standardu przenośnych baz wiedzy → przenośność + `/lint`), `source`, `summary` (jedna linia - to trafia do katalogu). 4. **Ujednolica język.** Baza trzyma jeden kanoniczny język (ten z onboardingu; przy konflikcie `CLAUDE.md`/Twoje ustawienia wygrywają nad domyślnym językiem skilla). Źródło w innym języku agent tłumaczy przy wsadzie; nazwy własne, kod, URL-e, daty i wikilinki zostają nietknięte. 5. **Rozstawia `[[wikilinki]]`** (linki między notami zapisywane w podwójnych nawiasach) do pokrewnych not i **dopisuje linki zwrotne** w tamtych notach - to zasila graf. 6. **Przenosi załączniki** (obrazki/PDF) do `content/ATTACHMENTS/` i poprawia odnośniki. 7. **Przenosi źródło** z `inbox/` do `_raw/processed/` (z datą). Pusty inbox = „zrobione”. 8. **Aktualizuje 3 indeksy**: `catalog.md` (linia noty), `vault-map.md` (licznik + tagi + „recent changes”), `graph.md` (linki wychodzące + linki zwrotne na celach). Jedno źródło potrafi dotknąć kilkunastu not - bo dokłada linki i linki zwrotne po całym grafie. ### Faza 3 - Kontrola (lista kontrolna) Na koniec agent sprawdza sam siebie: czy `total_notes` się zgadza, czy nowe noty są w „recent changes”, czy każda ma linię w `catalog.md`. Potem drukuje raport: ile źródeł, ile not powstało, ile scalono, stan indeksów (✅). Jak coś się nie zgadza - pokazuje rozbieżność, zanim powie „gotowe”. ## Co dokładnie się zmienia (diff) Po `/ingest` zajrzyj w Source Control - widzisz **na oczy, co się zmieniło**. Jedno źródło o Zettelkasten dotknęło **7 plików**: ![Diff po /ingest: nowa nota Zettelkasten.md, zaktualizowane indeksy i link zwrotny w innej nocie](/images/kurs/3-pierwszy-ingest-01.webp) - **Nowa nota** - `content/REFERENCE/Zettelkasten.md`: z frontmatterem, streszczeniem i wikilinkami. - **Zaktualizowana inna nota** - `Wikilinks Explained.md` dostał **link zwrotny**, bo nowa nota do niego linkuje (`[[Wikilinks Explained]]`), a graf jest dwukierunkowy. - **Trzy indeksy** - `catalog.md`, `graph.md`, `vault-map.md` odświeżone. - **Źródło** - `sample-source.md` zniknął z inboxu i wylądował w `_raw/processed/` (z datą). Nic nie kasujesz; jest odwracalne. Raport na końcu: _1 źródło → 1 nowa nota, 0 scalonych, indeksy ✅, inbox pusty._ Jeden wniosek: **pojedynczy plik wpiął się w graf i zaktualizował sąsiadów** - bez `/ingest` miałbyś tylko luźny plik w folderze. ## Surowe pliki vs wiki To różnica „**surowe pliki vs wiki**”: bez `/ingest` nie ma linków ani indeksów - masz tylko stertę plików. ## Pułapki - **Nie wrzucaj wszystkiego bez `/ingest`** - surowe pliki ≠ wiki (brak linków i indeksów). - **Nie edytuj indeksów ręcznie** - buduje je ingest (i `/reindex`, lekcja 4). Ręczna edycja się rozjedzie. - **Sprawdź `#todo/classification`** - jeśli agent nie był pewny folderu, oznaczy tak notę; przejrzyj ją. - **Inbox po `/ingest` ma być pusty** - jak coś zostało, raport powie dlaczego (np. nieudany URL). ## FAQ ### Czym `/ingest` różni się od zwykłego wrzucenia pliku do folderu? Wrzucony plik to luźny plik - bez frontmatteru, linków i wpisu w indeksie. `/ingest` klasyfikuje źródło, dobiera typ noty, wypełnia frontmatter OKF, rozstawia `[[wikilinki]]` i linki zwrotne, archiwizuje źródło i aktualizuje trzy indeksy. To różnica między stertą plików a żywą wiki. ### Co bierze `/ingest` - tylko wskazany plik? Bierze **wszystko, co leży w `content/_raw/inbox/`** - nie podajesz nazwy pliku. Możesz też podać link YouTube jako argument (`/ingest https://youtu.be/...`); agent ściągnie transkrypt (`yt-dlp`, w razie braku napisów awaryjnie przez Whisper/`ffmpeg`) i potraktuje go jak zwykłe źródło. ### Czy ingest nadpisze albo skasuje moje istniejące noty? Nie. Gdy w `catalog.md` znajdzie podobną notę, robi **scalenie** - dokłada treść, nie nadpisuje Twojej. Gdy podobnej nie ma, tworzy nową notę z szablonu. Źródło nie znika - wędruje z `inbox/` do `_raw/processed/` (z datą), więc wszystko jest odwracalne. ### W jakim języku powstaną noty, jeśli źródło jest po angielsku, a baza po polsku? Baza trzyma **jeden kanoniczny język** ustalony przy onboardingu, a przy konflikcie `CLAUDE.md` / Twoje ustawienia wygrywają nad domyślnym językiem skilla. Źródło w innym języku agent tłumaczy przy wsadzie. Nazwy własne, kod, URL-e, daty i wikilinki zostają nietknięte. ### Skąd wiem, że ingest zrobił wszystko poprawnie? Faza 3 to lista kontrolna: czy `total_notes` się zgadza, czy nowe noty są w „recent changes”, czy każda ma linię w `catalog.md`. Na koniec dostajesz raport (ile źródeł, ile not powstało, ile scalono, stan indeksów). Dodatkowo zajrzyj w Source Control - zobaczysz na oczy pełny diff. ### Jedno źródło zmieniło kilkanaście plików - czy to normalne? Tak. Nowa nota dokłada `[[wikilinki]]` do pokrewnych not, a do każdej z nich dopisuje **link zwrotny** - graf jest dwukierunkowy. Do tego dochodzą trzy indeksy i przeniesienie źródła. Dlatego jeden mały plik potrafi dotknąć kilkunastu not: właśnie wpina się w graf i aktualizuje sąsiadów. --- # Pytania i zarządzanie Source: https://pawel.lipowczan.pl/llm-wiki/kurs/4-pytania-i-zarzadzanie Cel tej lekcji: pytać bazę (nie czat) i utrzymywać jakość. Po lekcji umiesz `/qa`, `/lint`, `/reindex` i rozróżniasz `qa` vs `research`. Onboarding (L2) i ingest (L3) masz za sobą - baza żyje. Teraz ją **używasz**: pytasz i pilnujesz jakości. Wszystkie komendy siedzą w `.claude/commands/`. ![Komendy zarządzania bazą w .claude/commands/ i start /lint w terminalu](/images/kurs/4-pytania-i-zarzadzanie-01.webp) ## `/qa` - pytaj bazę, nie czat `/qa "twoje pytanie"` to sedno. Agent nie zgaduje - **czyta bazę** przez progressive disclosure (od ogółu do szczegółu): `vault-map` → `catalog` → `graph` → otwiera **tylko trafne noty** → syntetyzuje odpowiedź **z `[[cytowaniami]]`** i **oznacza luki**, gdzie baza milczy. Różnicę wobec czatu najlepiej widać, gdy pytasz o coś, **czego w bazie nie ma**. Zapytaj `/qa OKF` na bazie o samych notatkach o notowaniu: ![/qa o OKF - brak pokrycia: baza mówi „nie mam tego”, nie zmyśla](/images/kurs/4-pytania-i-zarzadzanie-05.webp) Baza sprawdza wszystkie noty, widzi **zero pokrycia** i mówi wprost: _„nie zmyślam odpowiedzi”_. Zamiast halucynacji dostajesz uczciwe „nie mam” + co dalej (`/ingest`, web research, `/gaps`). **Czat zgaduje; baza cytuje albo przyznaje, że nie wie.** Teraz to samo dla tematu, który **jest** w bazie - `/qa` o Zettelkasten: ![/qa o Zettelkasten - synteza z cytowaniami źródeł, oznaczone luki i follow-upy](/images/kurs/4-pytania-i-zarzadzanie-06.webp) Dostajesz syntezę z **cytowaniami** (`[[Zettelkasten]]`, `[[Wikilinks Explained]]`, `[[Example Note]]`), listę **luk** (czego bazie brakuje) i trzy follow-upy: zapis do `_outputs/answers/`, `/compile`, `/enhance`. Odpowiedź jest osadzona w **Twojej** wiedzy, nie w treningu modelu. ## `/compile` i `/enhance` - z odpowiedzi w trwałą notę `/qa` daje odpowiedź; te dwie komendy zamieniają ją w **trwały** kawałek bazy: - **`/compile `** - składa **nową notę zbiorczą / artykuł** z wielu istniejących not (synteza wyższego rzędu). Dobre, gdy materiał jest rozsypany po notach i chcesz go spiąć w jedno. - **`/enhance `** - **poprawia/rozbudowuje jedną notę**: dokłada sekcje, wikilinki, łata luki wskazane przez `/qa` lub `/lint`. W skrócie: `/compile` **tworzy** syntezę z wielu źródeł, `/enhance` **ulepsza** pojedynczą notę. ## `/lint` - przegląd stanu Baza po cichu gnije: martwe linki, sieroty, niespójne tagi, przeterminowane noty. `/lint` to audyt - sprawdza **10 klas problemów** i zapisuje raport do `_outputs/reports/`. Zaczyna, jak każda operacja, od indeksów (progressive disclosure - nie skanuje całości na ślepo): ![/lint startuje: czyta indeksy zgodnie z Navigation Protocol](/images/kurs/4-pytania-i-zarzadzanie-02.webp) Potem raportuje znaleziska po klasach - tu baza zdrowa, 3 drobne znaleziska: ![/lint - raport: 2 martwe linki, 1 brakujące połączenie, reszta czysta](/images/kurs/4-pytania-i-zarzadzanie-03.webp) Raport ląduje jako nota w `_outputs/reports/`, z tabelą wszystkich 10 klas i licznikiem: ![Raport zdrowia bazy: tabela 10 klas problemów z liczbami](/images/kurs/4-pytania-i-zarzadzanie-04.webp) Co sprawdza (10 klas): brakujący/niepełny frontmatter · zepsute wikilinki · sieroty (noty, do których nic nie linkuje) · stuby (niedokończone noty-szkice) · niespójne tagi · `#todo` · brak `summary` · brakujące połączenia semantyczne · przeterminowana treść (>1 rok bez przeglądu) · zgodność `type` z `CLAUDE.md`. Na koniec proponuje naprawę (tu: martwe linki → `/reindex`, brakujący link → `/enhance`). **Odpalaj co jakiś czas** - inaczej problemy się kumulują. ## `/reindex` - przebuduj indeksy Gdy indeksy się rozjadą (ręczna edycja, przerwany ingest), `/reindex` **buduje je od zera** z not: `vault-map.md`, `catalog.md`, `graph.md`. ![/reindex - deterministyczna przebudowa indeksów z wyrywkowym sprawdzeniem liczb](/images/kurs/4-pytania-i-zarzadzanie-07.webp) To **deterministyczna** przebudowa: odtwarza indeksy z aktualnego stanu not i sprawdza liczby (`total_notes`, węzły, krawędzie). Ważne: jeśli problem siedzi w **skrypcie** (np. parser policzył `[[...]]` ze środka bloku kodu jako realny link), reindex go **nie naprawi** - wiernie odtworzy ten sam wynik. Wtedy poprawka idzie do skryptu, nie do danych. Reindex leczy rozjazd indeks↔noty, nie błędy parsera. ## `qa` vs `research` - czemu osobno `/qa` odpowiada **z tego, co już masz**. Skille research **dokładają nową wiedzę z zewnątrz** i odkładają ją do bazy: `/research` (ukierunkowany), `/research-deep` (wieloźródłowy + weryfikacja), `/research-report` (research + raport), `/research-add-fields|items` (rozszerzanie list). Sedno: `qa` = czytasz bazę; `research` = baza rośnie. ## Pełna ściąga komend | Komenda | Do czego | | ------------ | ------------------------------------- | | `/onboard` | Konfiguracja bazy | | `/ingest` | Źródło → nota + indeksy | | `/qa` | Pytanie → synteza z cytowaniami | | `/compile` | Artykuł/nota zbiorcza z wielu źródeł | | `/enhance` | Popraw/rozbuduj notę | | `/lint` | Przegląd stanu / jakości | | `/reindex` | Przebuduj indeksy | | `/curate` | Wycofaj słabe/stare noty do `_graveyard/` | | `/gaps` | Znajdź luki / brakujące tematy | | `/refactor` | Przebuduj strukturę/noty | | `/output` | Wygeneruj raport/eksport | | `/research*` | Autonomiczny research | ## Zasady jakości - index-first / progressive disclosure (start od `vault-map`, nie skanuj całości), - frontmatter z `type` (OKF, przenośność), - aktualizuj indeksy po każdym zapisie, - zgodność z OKF (`git clone` → masz), - oczyść przed dzieleniem (usuń wrażliwe dane + atrybucje). ## Anty-wzorce Ręczne pisanie wiki · wrzucanie wszystkiego bez `/ingest` · olewanie `/lint` · skanowanie całego `content/` zamiast indeksów (pali tokeny). ## FAQ ### Czym `/qa` różni się od zwykłego zapytania czatu? Czat odpowiada z treningu modelu i zgaduje - bez gwarancji, że w ogóle zajrzał do Twoich not. Zapytany wprost agent sam decyduje, czy i jak szukać: potrafi wczytać za dużo (koszt tokenów) albo za mało (pominie właściwą notę), a wynik zmienia się między uruchomieniami. `/qa` wymusza za każdym razem ten sam protokół: **czyta Twoją bazę** (index-first: `vault-map` → `catalog` → `graph` → trafne noty), syntetyzuje odpowiedź z `[[cytowaniami]]` i oznacza luki, gdzie baza milczy. Dostajesz odpowiedź osadzoną w Twojej wiedzy, powtarzalną i możliwą do sprawdzenia w źródłach. ### Co, gdy zapytam `/qa` o coś, czego nie ma w bazie? Zamiast halucynować, baza raportuje **zero pokrycia** i mówi wprost: „nie zmyślam odpowiedzi”. Dostajesz uczciwe „nie mam tego” plus co dalej: `/ingest` nowego źródła, research z sieci albo `/gaps`. To kluczowa różnica: czat zgaduje, baza cytuje albo przyznaje, że nie wie. ### `/compile` vs `/enhance` - kiedy którego użyć? `/compile ` składa **nową notę zbiorczą / artykuł** z wielu istniejących not - synteza wyższego rzędu, gdy materiał jest rozsypany. `/enhance ` **poprawia jedną notę**: dokłada sekcje, wikilinki, łata luki wskazane przez `/qa` lub `/lint`. Skrótowo: compile tworzy z wielu źródeł, enhance ulepsza pojedynczą notę. ### Jak często odpalać `/lint`? Co jakiś czas - baza po cichu gnije (martwe linki, sieroty, niespójne tagi, przeterminowane noty), a problemy się kumulują. `/lint` sprawdza 10 klas problemów i zapisuje raport do `_outputs/reports/` wraz z propozycją naprawy. Traktuj to jak okresowy przegląd stanu. ### `/reindex` naprawi każdy problem z indeksami? Nie. To **deterministyczna** przebudowa - odtwarza `vault-map.md`, `catalog.md`, `graph.md` z aktualnego stanu not. Leczy rozjazd indeks↔noty (ręczna edycja, przerwany ingest), ale jeśli błąd siedzi w **skrypcie** (np. parser liczy `[[...]]` ze środka bloku kodu), reindex go nie naprawi - wiernie odtworzy ten sam wynik. Wtedy poprawka idzie do skryptu, nie do danych. ### Czym `/qa` różni się od skilli `research`? `/qa` odpowiada **z tego, co już masz** w bazie. Skille `research` **dokładają nową wiedzę z zewnątrz** i odkładają ją do bazy. Sedno: `qa` = czytasz bazę, `research` = baza rośnie. Więcej o research w kolejnej lekcji. --- # Rozwój i publikacja Source: https://pawel.lipowczan.pl/llm-wiki/kurs/5-rozwoj-i-publikacja Cel tej lekcji: domknąć zestaw komend, opublikować bazę i poznać ścieżkę rozwoju. Po lekcji znasz **pozostałe komendy** (porządki, generowanie, analiza), umiesz publikować przez Quartz (generator stron WWW z plików markdown), rozumiesz przenośność OKF (Open Knowledge Format - standard przenośnych baz wiedzy) i wiesz, dokąd baza rośnie dalej. ## Pełny zestaw - pozostałe komendy Rdzeń masz z lekcji 1-4: `/onboard`, `/ingest`, `/qa`, `/lint`, `/reindex`, plus `/compile` i `/enhance`. Zostały komendy, po które sięgasz rzadziej - do porządków, generowania i analizy. Wszystkie trzymają się tej samej zasady: **czytaj indeksy, aktualizuj indeksy, nie kasuj bez potwierdzenia.** ### `/gaps` - co zbudować dalej Analiza **luk wiedzy** - nie mechanicznych problemów (tym zajmuje się `/lint`), tylko braków merytorycznych: noty słabo połączone (stopień linków ≤ 1), tematy implikowane, lecz nieopisane, cienkie foldery i tagi, przeterminowane klastry. Zwraca uszeregowaną listę „do zbudowania / do połączenia” i proponuje zapis do `content/_outputs/reports/`. Nie zmienia bazy - to kompas, nie koparka. Odpalaj, gdy nie wiesz, co notować następne. ### `/curate` - sprzątanie (przeterminowane noty) Okresowa higiena: ocenia każdą notę (wiek, izolacja, martwe linki, duplikaty), pisze raport segregacji do `content/_outputs/reports/` i **dopiero po Twoim potwierdzeniu** wycofuje przeterminowane noty do `content/_graveyard/`. Motto: _„`/lint` diagnozuje; `/curate` leczy”_. Domyślnie to próbny przebieg - sam raport. Akcje na notę: `archive` (→ `_graveyard/`), `merge` (→ `/refactor`), `refresh` (→ `/enhance`), `keep`. Wycofanie jest **odwracalne** (przeniesienie, nigdy `git rm`; `_graveyard/` jest wykluczony z indeksów i publikacji). Kadencja: raz na kwartał. ### `/refactor` - przebuduj bez psucia linków Rename / move / merge / split not z automatyczną naprawą **wszystkich** `[[wikilinków]]` i przebudową indeksów. Linki rozwiązują się po nazwie pliku, więc ręczne przenoszenie je psuje - `/refactor` robi to bezpiecznie (zachowuje aliasy `[[A|B]]` i kotwice `[[A#nagłówek]]`). Po każdej operacji odpala reindex i proponuje `/lint`. Używaj do zmian strukturalnych, których nie chcesz robić ręcznie. ### `/output` - artefakty pochodne Generuje **pochodny** artefakt z bazy: podsumowanie, listę lektur, mapę tematu albo oś czasu. Domyślnie zapisuje do `content/_outputs//`. Różnica wobec `/compile`: `/compile` tworzy pełny **artykuł** (notę `compiled-note` w folderze tematu - wchodzi do wiki i indeksów), a `/output` daje raport/zestawienie **obok** wiki. Krótko: compile publikuje wiedzę, output ją eksportuje. ### Rodzina `research` - przedsmak (osobna lekcja) `/research`, `/research-add-items`, `/research-add-fields`, `/research-deep`, `/research-report` to potok, który **dokłada nową wiedzę z zewnątrz**. Buduje szkielet badania (`outline.yaml` + `fields.yaml`) w `content/_raw/research-workspaces/`, odpala po jednym agencie na element (ustrukturyzowany JSON per element), a na końcu składa raport i wrzuca go do `content/_raw/inbox/` - gotowy pod `/ingest`. To dzięki niemu baza rośnie sama. Rozbijemy go w **osobnej lekcji** - tu tylko sygnalizuję, że istnieje. ### `excalidraw-diagram` (skill) - diagramy Skill generujący diagramy `.excalidraw` (proces, architektura, koncept) z pętlą samo-walidacji przez Playwright. Wymaga jednorazowej konfiguracji (`uv` + Playwright/Chromium). Opcjonalny dodatek, gdy chcesz do noty wrzucić schemat, a nie tylko tekst. ## Publikacja (opcja) `npx quartz build` → strona WWW (jak `brain.lipowczan.pl`). Markdown zostaje markdownem; Quartz to tylko warstwa publikacji. ## OKF / przenośność Markdown + frontmatter + `index.md`/`log.md` = baza, którą da się wymienić. „If you can `git clone` it, you can ship it.” ## Ścieżka rozwoju - Podłącz brain do **systemu agentowego** (multi-brain): jeden agent odpytuje wiele baz (wzorzec `brain-query`). - **MCP** (protokół, którym narzędzia AI podłączają się do zewnętrznych źródeł; tu: `brain-mcp`): udostępnij bazę dowolnemu klientowi (Claude Desktop / IDE). - **Publikacja Quartz** → marka osobista / portfolio wiedzy. - **Wymiana paczek OKF** - eksport dopracowanej bazy jako produkt (ekstrakt: sama wiedza, bez plików roboczych). ## 🤖 Gotowe prompty - co dalej Znajdź, czego bazie brakuje, i przygotuj publikację. Wklej na **swoim repo bazy**: ```text Odpal skill gaps: znajdź luki wiedzy - słabo połączone noty, brakujące tematy, cienkie obszary. Zaproponuj 5 konkretnych źródeł lub pytań, którymi domknę luki. ``` ```text Zbuduj publiczną wersję bazy przez Quartz (npx quartz build) i wypisz krok po kroku, jak wypchnąć ją na GitHub Pages. Nie publikuj niczego bez mojej zgody. ``` ## FAQ ### Czym jest OKF i dlaczego baza jest przenośna? OKF (Open Knowledge Format) to standard opisany przez Google, który formalizuje wzorzec „wiedza jako zwykły markdown z frontmatterem”. Twoja baza to zwykłe pliki `.md` z polem `type` - czytelne i dla agenta, i dla człowieka, bez uzależnienia od konkretnej aplikacji. Stąd „jeśli możesz to `git clone`, możesz to wysłać”: przenosisz repo i cała wiedza działa dalej. ### Czy muszę publikować bazę? Nie. Publikacja przez Quartz jest w pełni opcjonalna - baza świetnie działa jako **prywatne** repo markdown, którego używasz tylko przez Claude Code. Quartz to warstwa na wierzchu; dokładasz ją tylko, jeśli chcesz mieć publiczną stronę wiedzy. ### Czy publikacja Quartz ujawni moje prywatne noty? Publikujesz tylko to, co sam zbudujesz i wypchniesz - Ty kontrolujesz zakres. Foldery robocze (`_raw/`) i wycofane noty (`_graveyard/`) są wykluczone, a przed dzieleniem usuwasz wrażliwe treści i atrybucje. Nic nie trafia na zewnątrz bez Twojej wyraźnej zgody. ### Czym różni się `/gaps` od `/lint`? `/lint` szuka **mechanicznych** problemów: martwe linki, brakujący frontmatter, sieroty, niespójne tagi - 10 klas higieny. `/gaps` szuka **luk wiedzy**: słabo połączonych not, brakujących tematów, cienkich obszarów. Skrótowo: `/lint` to higiena, `/gaps` to strategia - co napisać następne. ### Czy `/curate` skasuje moje noty? Nie kasuje. Domyślnie robi próbny przebieg - sam raport segregacji. Przeterminowane noty wycofuje do `content/_graveyard/` **dopiero po Twoim potwierdzeniu**, i to odwracalnie (przeniesienie, nigdy `git rm`). Cofnięcie = przeniesienie noty z `_graveyard/` z powrotem. ### Kiedy użyć `/output`, a kiedy `/compile`? `/compile` tworzy pełny **artykuł** - nową notę `compiled-note` w folderze tematu, która wchodzi do wiki i indeksów. `/output` generuje **pochodny** artefakt obok wiki (podsumowanie, lista lektur, oś czasu) do `content/_outputs/`. Chcesz trwały wpis w bazie → compile; chcesz zestawienie na wynos → output. ## Most do wersji płatnej To był pusty szablon. Pełniejsze akceleratory - gotowe, sprawdzone skille i **ekstrakt realnej bazy dopracowanej miesiącami** (gotowe paczki wiedzy do załadowania do swojej bazy) - szykuję jako płatne paczki. Pomijasz tygodnie iteracji i tysiące spalonych tokenów. Repo szablonu: [github.com/plipowczan/second-brain-template](https://github.com/plipowczan/second-brain-template).