Can AI Truly Take Ownership of Code?
We set out to build an AI agent capable of acting as a full-fledged software developer — one that could support the real-world migration of Actualog’s PIM solution from .NET 4.6 to .NET 9.0 MVC Core. But we weren’t aiming for a tool that merely generates a C# class or a JavaScript module. Our vision was an agent that takes true code ownership — one that understands Actualog’s design patterns, internal architecture, and can work independently. No more babysitting. We’re not looking for a clever junior developer; we’re engineering a reliable, self-directed teammate. This AI should be able to plan, implement, and evolve complex features while maintaining an active dialogue with a product owner or business analyst. It must do more than write code — it needs to think like an architect, apply patterns intentionally, enforce systemic consistency, and make smart use of what’s already been built.
Knowledge Transfer
As we build our RAG-augmented AI developer to support the migration of Actualog’s legacy solution, we’re also learning how to architect, fine-tune, and operationalize intelligent agents in complex software environments. These learnings are not confined to the migration task. The real strategic payoff is in transferring this knowledge — the retrieval workflows, architecture-aware prompting, validation techniques, and knowledge graph integration — into Actualog PIM itself.
Every retrieval strategy, every design pattern encoded into the agent, and every insight about AI-human collaboration becomes part of the foundation for a more intelligent, maintainable, and future-proof Actualog platform.
The Cognitive Trinity
To function at the level of an autonomous developer, the AI must learn to reason. Not just statistically, but structurally and semantically. This learning is scaffolded through three interdependent sources:
- Vector Database: Enables the retrieval of semantically similar code, documentation, and usage examples via dense embeddings — grounding code generation in prior art.
- Knowledge Graph: Maps structural and functional relationships among components, layers, and flows — supporting architectural coherence and context awareness.
- Design Patterns & Architecture Blueprints: Encodes domain-specific architectural wisdom, promoting consistency and enforcing design conventions during generation.

Together, these modalities form a multi-perspective reasoning system. This goes beyond traditional code synthesis, enabling Retrieval-Augmented Generation (RAG) for architecture-aware software creation — where the AI isn’t just generating snippets, but thinking in terms of reusable patterns, systemic behavior, and long-term maintainability. Below is the research by ChatGPT o3.
Introduction
Building an AI-assisted developer that can own an entire .NET 9.0 codebase requires more than just code generation. Such a system must deeply understand the project’s architecture, design patterns, and coding standards. The envisioned setup is an AI developer agent working alongside a human project manager (PM) or business analyst, capable of implementing features and refactoring code autonomously. To achieve this, the AI must leverage Retrieval-Augmented Generation (RAG) – injecting relevant project knowledge into its prompts – so that it produces code aligned with existing patterns and best practices. This research report explores the most efficient architecture for such a system, combining LLM capabilities with rich context from the codebase to enforce consistency and correctness. We survey academic insights, industry case studies, and practical tools that inform the design of an AI developer agent for a complex .NET project.
Challenges and Requirements
Designing a “sole developer” AI agent entails several challenges unique to software engineering:
Understanding Complex Architecture: The AI must integrate and reason about the project’s architecture and design patterns. Current code LLMs often struggle to follow project-specific conventions or patterns, producing code that conflicts with required designs arxiv.org. The agent needs a way to absorb the project’s domain knowledge – e.g. whether the project uses MVC, microservices, CQRS, etc. – so it can adhere to those structures in new code.
Enforcing Consistency and Best Practices: All generated code should comply with established layer boundaries (controllers, services, models, views), reuse existing components where possible, and conform to naming and style guidelines. Without explicit guidance, LLMs may ignore subtle architectural rules (e.g. “controllers must never call the database directly”) leading to architecture drift. Ensuring consistency requires the AI to be aware of architecture rules that have been defined for the codebase. Tools exist for human developers (like ArchUnit for Java or NetArchTest for .NET) to encode such rules infoworld.com, and similar checks or knowledge must be embedded into the AI’s workflow. If architecture rules are not regularly validated, a codebase’s design will degrade over time. The AI agent should effectively act as a guardian of these rules.
- Full Codebase Context: The codebase is likely too large to fit entirely into the LLM’s prompt window at once. Yet, the agent needs relevant context (APIs, data models, existing function usage) to answer questions or generate new code correctly. Without repository-specific knowledge, an LLM can only give generic output that often hallucinates or misuses APIs ar5iv.org. Retrieval-Augmented Generation addresses this by fetching relevant code snippets or documentation and supplying them in-context. The system must efficiently index and search the codebase so that, for any given task, the LLM sees the most pertinent pieces (e.g. a similar method to mimic, or a pattern description) sourcegraph.com.
- Embedded Domain Knowledge: Beyond raw code, the agent should leverage architectural documentation, design decision records, and pattern templates specific to the project. For example, if the project has a standard “error handling pattern” or a JSON schema for API responses, that knowledge should be retrievable. The AI’s knowledge base may include not only code text, but also graph metadata (like a call graph or module dependency graph) and inferred links between components (e.g. which UI component corresponds to which backend service). This ensures the AI can reason over relationships (such as “Feature X touches A, B, and C components”) even if such links aren’t obvious from isolated code chunks. Advanced approaches propose building a code knowledge graph capturing functions, classes, and their relationships (calls, inheritance, data flows) daytona.io. With a knowledge graph, an AI assistant can answer higher-level questions like “what’s the most complex part of the auth system?” by traversing connected context, not just doing keyword search.
- Human-in-the-Loop and Decision Making: The AI developer should not operate blindly; it will work with a human PM for guidance. This means the system should support partial human approval cycles. For instance, the AI might propose two different implementation approaches for a feature (trading off complexity vs. performance) and allow the PM to choose. The agent needs to present multiple solution paths when appropriate, explain the implications of each, and possibly adjust its plan based on feedback. This requires an element of reasoning and dialogue in the agent: it must justify how its suggestion aligns with the project’s architecture (“Approach A uses our existing caching module as per our pattern, while Approach B introduces a new dependency which might violate our microservice isolation”).
With these requirements in mind, we now outline an architecture that can meet them, followed by examples and existing systems that inform each aspect.
High-Level System Architecture
At a high level, the RAG-powered developer agent consists of several coordinated components, each addressing a piece of the challenge:
- Codebase Knowledge Ingestion: A pipeline to embed and index the project’s knowledge. All source code, documentation, and relevant design artifacts are processed here. For source code, the system can parse it into logical chunks (e.g. one class or one method per chunk) and generate vector embeddings that capture semantic meaning. These embeddings are stored in a vector database for fast similarity search medium.com. Alongside the vector index, a symbol index (for exact symbol lookups) and potentially a keyword index can be maintained for precise matching (a hybrid dense/sparse search approach). This mirrors how Sourcegraph’s Cody uses both semantic embeddings and traditional keyword search to find code context, yielding better results than vector search alone latent.space. The ingestion step may also extract an architecture graph: using static analysis to record relations like “Class A calls Class B” or “Module X depends on Module Y”. The result is a rich knowledge base combining text embeddings, graph relationships, and metadata.
- Pattern & Constraint Knowledge Base: In parallel to raw code, the system needs a structured representation of the project’s architectural patterns and rules. This can be a collection of JSON/YAML definitions, a knowledge graph, or even unit tests that assert architectural constraints. For example, a pattern document might declare how the Repository pattern is implemented in this project (which base class all repositories inherit, how they should be used), or define layering rules (e.g. “UI layer can call only Service layer, not Data layer directly”). These rules could be encoded as graph constraints (nodes representing components with
layerattributes, and disallowed edges between certain layers) or in a DSL that the AI can interpret. By encoding design rules explicitly, the agent can both validate generated code and use the rules to guide generation (for instance, if asked to create a new data access class, it knows it must conform to the IRepository interface, etc.). Some organizations use tools like ArchUnit/NArchTest to formalize these rules in tests infoworld.com; an AI agent could query such tests to see what is expected. In essence, this knowledge base is the AI’s internal guide to “what is allowed or recommended in this codebase.” - Retrieval Engine: When the AI is given a task or question, the retrieval component kicks in to gather relevant context from the knowledge bases. This engine likely implements a two-stage retrieval: first, a broad recall from the vector index (finding many code snippets or docs that might relate to the query), then a ranking or filtering stage to pick the most relevant pieces sourcegraph.com. For example, if the query is “Add a new payment API endpoint,” the retrieval might pull (a) the existing API controller template or a similar endpoint, (b) any design doc about payment integration, (c) the service class and data model related to payments, and (d) the coding standard for controllers. These snippets are then provided as context to the LLM. The retrieval logic can also use the graph: e.g. find the component named “PaymentService” in the knowledge graph, get its related nodes (methods, models), and retrieve those code snippets for context. By combining semantic search with graph-based lookup, the agent ensures it retrieves not just textually similar code, but architecturally relevant code (e.g. all usage sites of a certain pattern). This context assembly is critical – as Jan Hartman et al. note, providing repository-specific snippets to the LLM dramatically increases the quality and accuracy of responses, essentially giving the model a “mini reference manual” of the codebase ar5iv.org.
- AI Reasoning and Generation (LLM Core): At the core is the Large Language Model (such as GPT-4, Code Llama, etc.) which generates answers or code using the retrieved context. The LLM is prompted with a structured prompt containing (a) the user’s request (e.g. feature description or question), (b) the selected context snippets (code, docs), and (c) instructions or few-shot examples enforcing style. Modern coding assistants use prompt techniques like Fill-in-the-Middle for code generation (so the model sees code before and after the insertion point) github.blog. The LLM in this system would be instructed to strictly follow the project’s patterns: for instance, the prompt might say “You are an expert .NET developer following the company’s internal architecture guide. Use only the existing utilities for logging,” etc. If the knowledge base has structured patterns, the prompt can include a summary of relevant rules (e.g. “NOTE: All database access must go through the
PaymentRepositoryclass.”). The model then produces output: which could be a code diff, a new file content, or an explanation. Importantly, the LLM should be capable of not just single-turn generation but multi-turn reasoning. That is, the agent might internally break a complex task into sub-steps: understand request -> plan solution -> retrieve specific contexts -> generate code for part A -> generate code for part B -> integrate -> review. Frameworks like LangChain or Semantic Kernel can orchestrate such multi-step prompts. In fact, this design aligns with the emerging “agentic” AI patterns, where an AI can plan and call tools iteratively rather than one-shot answers. - Verification and Testing: After generation, the system should verify that the new code indeed compiles, passes tests, and adheres to the patterns. This is akin to having an integrated CI/CD in the loop. The agent can invoke a code compiler and test runner (e.g. use Roslyn to compile the C# code, run unit tests, etc.) as tools. It can also run static analysis or the aforementioned architecture tests to catch any rule violations. For example, if the AI accidentally introduced a dependency rule violation, a NetArchTest unit test would fail – the agent could detect that and adjust the code. This feedback loop is crucial for an autonomous coding agent: it’s not enough to produce code, it must ensure the code integrates well. GitHub’s latest Copilot “agent” mode hints at this capability by auto-fixing errors in a loop github.blog. By checking its work, the AI can move closer to a zero-regression policy, where human intervention is only needed for truly novel design decisions or ambiguous requirements.
- Human Interface (Approval Workflow): Finally, the system presents its work to the human project manager or reviewer. The AI should generate not only code, but also a rationale: describing what it did, which existing patterns it used, and any alternatives considered. If multiple solution paths were explored, the AI can summarize them. For instance: “I implemented the payment API by reusing our existing
PaymentServicelogic. Alternatively, I considered creating a new service, but that would duplicate functionality, so I chose the former for consistency.” This gives the human confidence that the AI’s changes are aligned with the overall architecture. The code changes might be delivered as a pull request for review, with the AI as the author and the PM as the approver. During review, the PM could ask follow-up questions (“Why didn’t you use caching here?”) and the AI, thanks to RAG, can answer by citing relevant parts of the code or documentation (“Because according to our architecture guidelines, the PaymentService handles caching internally ar5iv.org, the controller should not implement caching.”). Once approved, the code merges, and the knowledge base can be updated if needed (e.g. new patterns learned).
In summary, the architecture is an orchestration of retrieval and generation: the context engine finds the right pieces of knowledge, the LLM core produces solutions, and a validation layer ensures compliance. This modular design allows optimizing each part (for example, improving the retrieval with better embeddings or fine-tuning the LLM on .NET code for higher accuracy).
Integrating Architectural Knowledge and Patterns
A distinguishing feature of this system is the emphasis on architectural pattern knowledge. Traditional code assistants mostly focus on immediate context (a few files) and general programming knowledge. Here, we need the agent to be explicitly aware of the architecture as a first-class entity. Some concrete strategies to encode and leverage this knowledge include:
- Knowledge Graph of the Codebase: As mentioned, representing the codebase as a graph can significantly enhance context understanding. Nodes can represent classes, interfaces, modules, and even higher-level concepts (like “feature” or “layer”), while edges capture relationships (calls, inherits, uses, belongs-to-layer). Enterprises like Deutsche Telekom found that using knowledge graphs combined with RAG improved the accuracy of their AI coding assistant, allowing it to surface nuanced context specific to their codebase medium.com. A knowledge graph provides a structured backdrop that pure text embeddings lack – it ensures the AI respects relationships that might not be obvious from isolated text. For example, a graph can tell the agent that
OrderControlleris part of the “Web” layer and should only communicate with services in the “Business Logic” layer, not directly with “Data Access” classes. When generating code, the agent could query the graph to see where a new class should fit or what existing components are related. This approach echoes the idea of giving the codebase a “brain” that the AI can consult for non-local knowledge daytona.io. Modern techniques even leverage LLMs to help build these graphs (parsing code and documentation to extract relationships) daytona.io. - Embedded Pattern Templates and Examples: The system should include a library of pattern examples or templates gleaned from the codebase. For instance, if the project uses a Repository Pattern for data access, the knowledge base should contain one canonical example of a repository class and how it’s used by a service. This can be used as a few-shot example when the AI needs to generate a new repository – the prompt can say “Here is how we typically implement a repository” followed by the example code. Academic research has begun looking at whether code LLMs truly understand design patterns or need help; one study found that without guidance, LLMs often miss project-specific design patterns and require developers to adapt the output to the project’s design manually arxiv.org. By feeding pattern examples into the prompt (a form of in-context learning), we guide the model to produce code that fits those patterns. Over time, as the model generates code and it’s validated, we could even fine-tune a specialized model on the project’s code (though RAG alone might suffice if done well).
- Constraints and Validation as Feedback: As part of generation or post-generation, the AI can use the encoded constraints (from the pattern knowledge base) to self-check its outputs. For example, if there’s a rule that all SQL queries must go through a certain utility, the agent could scan its generated code for any raw SQL usage and flag it. This is analogous to how linters or static analyzers enforce rules. The difference is the AI can then fix any violations on the fly. In essence, the patterns and rules are not only reference material but also act as tests that the AI’s output must pass (either via actual unit tests like NetArchTest or via internal checks). This creates a robust loop where the AI “thinks with” the pattern knowledge. We can draw inspiration from tools like Apiiro’s LLM-based pattern detection, which combines LLMs with pattern rules to identify deviations in code security patterns apiiro.com. Our agent similarly combines pattern knowledge with code generation to ensure alignment.
- Architectural Reasoning in Prompts: The prompts themselves can encourage the model to reason about architecture. Instead of only asking “Generate code to do X”, the system might prompt the model to first output a brief plan that names which layers or components will be involved. For example: “Plan: To implement feature X, I will add a method in YService, call it from XController, and update ZModel. This follows the existing 3-layer architecture.” Only after this plan (which the system can verify or adjust) does it proceed to code. This chain-of-thought prompting helps the model avoid tunnel-vision on coding and keep the big picture in mind. It’s analogous to a senior developer first figuring out how a feature fits in the system, then writing the code. Encouraging the model to explicitly mention architecture in its reasoning can reduce mistakes and make it easier for humans to follow the AI’s thought process.
In practice, integrating these elements means our AI agent is not a black-box code generator but a semi-expert system that mixes declarative knowledge (architecture rules, graphs) with the generative power of LLMs. The use of structured knowledge (graphs, JSON rules) in RAG is a growing trend; researchers call this approach Graph-RAG or knowledge-grounded code generation medium.com. It helps bridge the gap between unstructured neural output and the structured nature of software architecture.
Example Workflow Scenario
To illustrate how the system might operate, consider a concrete example scenario:
AI Agent Process:
- Understanding the request: The agent parses the request and identifies that this involves the authentication/registration subsystem. It queries the knowledge graph for “registration” or related components. It finds a
UserController(responsible for registration), aUserService, and anEmailService(which perhaps exists for other email notifications). It also retrieves an architecture rule: “All email sending is done via EmailService; controllers should not send emails directly.” - Retrieval of Context: Using the above info, the agent retrieves the
Register()method code fromUserController, theEmailServiceinterface and implementation, and any configuration related to email (perhaps a template or a setting for sender address). It also finds a pattern document about “Post-registration actions” if one exists, or similar feature code as example (maybe password reset confirmation email code). These snippets and notes are assembled as context. - Plan Formation: The agent forms a plan in natural language (this could be internal, or visible to the human in an interactive mode): e.g. “To send a confirmation email, I will modify UserController.Register to call EmailService.SendWelcomeEmail after creating the user. I will create a new template for the welcome email if not exists. This aligns with our pattern of keeping email logic in the service layer.” The PM reviews this plan and approves.
- Code Generation: Given the green light, the agent writes the code. It might create a new method
EmailService.SendWelcomeEmail(User user)if needed (reusing the EmailService infrastructure). It ensures to use the logging and error handling conventions as seen in similar methods (thanks to including an example in the context). The controller change is made respecting transaction boundaries or other concerns the context shows. The AI then produces a diff or the modified/new files. - Self-Verification: The agent runs the test suite. Suppose a unit test fails because it expected
Register()not to send emails (perhaps the test needs updating). The AI notices this and updates the test or marks that an expected behavior changed, providing an explanation. It also runs the architecture tests: all pass, since it usedEmailServiceas required (if it had directly used SMTP client in the controller, an ArchTest would flag a violation of layering). The agent also double-checks: did it follow the email template pattern? If the style guide says emails must be sent asynchronously, it verifies it used the async method of EmailService (adjusting code if not). - Output and Alternatives: The agent prepares the final output for the human: the code changes and a summary. It might say:“✅ Implemented email confirmation via
EmailService. The registration flow now callsSendWelcomeEmailafter user creation. All new code follows the existing patterns (see PasswordResetEmail for a similar implementation which I referenced).” If there was an alternative (like maybe using an event-driven approach), the agent could note: “I considered using our event bus to handle the email, which would decouple the controller, but since the system doesn’t currently have an event for user registration, I chose the direct service call for consistency. We can refactor to an event-based approach later if desired.” This demonstrates architectural reasoning. The PM can then merge the change or discuss the alternative if they prefer the event-driven path, in which case the AI can implement that instead.
User Story: “As a user, I want to receive an email confirmation after I register, so that I know my account was created.” The PM asks the AI to implement email confirmation in the existing registration flow of a .NET 9.0 web application.
Throughout this flow, we see the agent leveraging RAG at multiple points: retrieving code, referencing pattern knowledge, and even retrieving test expectations. It behaves akin to a diligent human developer who reads the docs and code before coding, writes the code, tests it, and documents the reasoning.
Case Studies and Existing Systems
This vision of an AI developer agent is ambitious but builds upon developments and experiments in both industry and academia. Here we highlight some real-world systems and research that inform our design:
- GitHub Copilot (OpenAI Codex): Copilot paved the way for code generation in IDEs, but originally it worked mainly by predicting code from the immediate context (open files, the current file’s content, etc.). GitHub has since introduced Copilot Chat and is exploring Copilot Extensions (Agents) that allow tools and retrieval. For example, Copilot can now use “neighboring tabs” to include relevant code from other files github.blog. GitHub has discussed RAG in the context of Copilot for Enterprises – by indexing a company’s private repositories and documentation, Copilot can retrieve relevant snippets as context. The GitHub Copilot Chat interface can answer questions about your code by effectively performing searches through the repository. While details of its retrieval algorithms are proprietary, the necessity is clear: as Idan Gazit of GitHub Research noted, “Without context, an LLM can only provide generic responses. With proper context, it can understand and reason about your specific code, architecture, and practices.” sourcegraph.com This principle is exactly what our system leverages. GitHub’s move to allow custom Copilot agents suggests they foresee specialized workflows – one could imagine a Copilot agent that enforces .NET architecture rules by consulting a company’s guidelines.
- Sourcegraph Cody: Sourcegraph’s Cody is explicitly designed to “know” your entire codebase. It uses Sourcegraph’s code search under the hood to fetch context for the LLM. Their published research (RecSys 2024 industry paper) highlights that providing the right context is key to relevant answers ar5iv.org. Cody’s architecture includes a context retrieval service which does hybrid search (combining dense vectors with precise token matching) latent.space. They mention achieving an industry-best code completion acceptance rate of 30% by using a context-enhanced LLM. Notably, Cody integrates with multiple LLMs (including open-source ones like StarCoder) and still reaches high quality by virtue of superior retrieval and prompt assembly latent.space. The “Normsky architecture” described by Sourcegraph’s team combines symbolic techniques (parsing, static analysis – a la Norvig) with LLMs (a la Chomsky) to ensure the AI’s suggestions are grounded latent.space. This validates our approach of mixing knowledge-based methods with neural generation. The lesson from Cody is that investing in code indexing and search infrastructure pays off significantly in making an AI assistant useful. Our system would similarly use advanced search to gather code context, and could even use Sourcegraph’s open APIs or similar tools for that purpose.
- Internal Developer Tools (Meta, Google, etc.): Big tech companies have been developing AI coding aids for their internal codebases. Meta’s CodeCompose (described in a 2024 paper) is an AI code completion tool deployed to 10,000+ developers arxiv.org. Meta fine-tuned their model on internal code to improve accuracy on their frameworks. Interestingly, they noted that including more contextual information (like the code after the cursor, file paths, etc.) significantly improved suggestion relevance. This again underscores the value of rich context. While CodeCompose mainly auto-completes code, Meta plans to incorporate “dynamic context from other files” to further improve it – effectively a retrieval mechanism across the repo. Google similarly has an internal system (often referenced as an AI pair programmer for Google’s monolithic codebase) that likely uses the company’s extensive code graph (Google’s code search, Kythe, etc.) to provide context. These cases show that at scale (millions of lines, many languages), retrieval and integration with existing code knowledge is essential. Our .NET AI developer might not operate at Google-scale, but even a few hundred thousand lines across services is non-trivial for an LLM to handle without retrieval.
- Academic Research on Code LLMs: Researchers are actively examining how well code LLMs comprehend software engineering constructs. The paper “Do Code LLMs Understand Design Patterns?” (Pan et al., 2025) found that out-of-the-box models often fail to recognize or correctly apply classic design patterns and project-specific styles arxiv.org. This leads to the need for manual post-processing by developers to make AI-generated code fit their codebase. Our approach directly addresses this gap by feeding pattern knowledge and enforcing it, essentially teaching the LLM about the project’s design at query time. Another relevant area is RAG with knowledge graphs – while much of the literature is on text QA, the concept of Graph-RAG (combining vectors with a knowledge graph) is emerging as highly effective for complex domains medium.com. In code, this could be revolutionary: imagine querying not just “similar code” but “code that follows the same sequence of operations” via graph matching. Early products like Cntxt (by Brandon Docusen, (open source) are demonstrating this: CntxtJS/JV analyzes codebases and produces an “optimized knowledge graph” of the code, which can be given to an LLM to drastically reduce the tokens needed to understand the project linkedin.com. It maps component relationships and outputs a summary that acts like “your codebase’s elevator pitch to the LLM”. The claim is up to 75% reduction in context size by using such graphs, while preserving understanding. This is a promising direction – our system could generate a similar summary or graph of the .NET project and always include it in prompts to keep the model grounded.
- Other AI Developer Agents: Beyond coding assistants that focus on suggestions, there are experimental agents aiming to autonomously build software. For example, projects like Replit’s Ghostwriter and OpenAI’s prototypes (as hinted in DevDay demos) show agents that can create whole projects from scratch via conversation. These often use iterative planning: e.g. writing code, executing it, debugging errors, etc., in a loop. While those are more general (and often start without an existing codebase), the patterns of iterative refinement and tool use are applicable to our case. Our agent should similarly be able to run code or tests as a “tool” and rectify issues. In fact, OpenAI’s DevDay demo included an AI agent that could use a terminal to run a app and fix mistakes – showcasing how tool integration allows the AI to go from writing code to verifying it in a runtime environment. Tools like Microsoft’s Semantic Kernel provide frameworks to define such iterative plans in .NET, which could be leveraged to implement the agent’s workflow (we can define skills for searching code, running build/test, etc., that the kernel orchestrates with an LLM).
In summary, the concept of a RAG-augmented AI developer is at the frontier of what companies are trying. GitHub and Sourcegraph are adding retrieval and agent capabilities to their copilots; enterprises are starting to combine knowledge graphs with LLMs for coding; and research is identifying the importance of pattern-awareness. Our proposed system stands on the shoulders of these developments, synthesizing them into a unified architecture tailored to .NET projects. Practical Implementation Guidance
Practical Implementation Guidance
To build and deploy such an AI developer system in a real enterprise setting, one should consider the following practical steps and best practices:
- Start with Read-Only Analysis: Initially, the AI can be introduced as a documentation aide or reviewer rather than letting it commit code autonomously. For example, use the RAG setup to answer developer questions about the codebase (“Which classes handle payment processing?”) or to review pull requests for architecture compliance. This lets you iteratively refine the knowledge base (embeddings, graph, rules) with lower risk. It also helps win trust from the team as they see useful, correct outputs.
- Construct the Code Index and Graph: Use existing tools to accelerate building the knowledge base. For .NET, you can use Roslyn analyzers or Mono.Cecil to parse assemblies and source code to extract a call graph, dependency graph, etc. Feed this into a graph database (Neo4j, Memgraph, etc.) or even a simple in-memory graph structure for quick lookup. For embedding, consider specialized code embedding models (like CodeBERT, UniXcoder, or OpenAI’s text-embedding-ada with code). Ensure to chunk code in logical units and include semantic info (function name, class name in the embedding metadata) to enable filtered searches (e.g. restrict search to “controller” layer by pre-labeling chunks).
- Integrate Vector Search with Symbol Search: A practical tip from Sourcegraph’s experience – pure vector search might retrieve something vaguely similar but not actually related to the specific identifier or module in question. Incorporate a way to do precise lookup. For example, if the user’s prompt or the AI’s plan mentions
EmailService, have the retrieval step explicitly pull up theEmailServicedefinition and references using a traditional search or an IDE index. This can be done by interfacing with something like Language Server Protocol (LSP) queries or a tool like Sourcegraph’s search API. Marrying these results with embedding-based results yields a more complete context. - Use Latest .NET Documentation: Since the agent works in .NET 9.0, make sure it has access to the latest API docs and framework source (if available). .NET evolves quickly; if the LLM was trained on .NET 6 or 7 code, it might not know new APIs in .NET 9.0. By adding official docs to the retrieval corpus, the agent can pull in usage examples or specs for any newer API it needs to use. For instance, if .NET 9.0 introduced a new library for email sending, retrieving its documentation ensures the AI uses it correctly rather than guessing. This addresses the “knowledge cutoff” problem – RAG gives the model “new” facts on demand medium.com.
- Model Selection and Fine-tuning: Depending on confidentiality and latency requirements, you might use an open-source code model (like Code Llama 34B, hosted internally) or an API like Azure OpenAI’s GPT-4. Fine-tuning the model on your codebase is an option, but it can be expensive and requires careful curation. Many organizations find RAG alone sufficient, and it avoids the need to retrain for every code change github.blog. That said, fine-tuning on a small set of “style exemplars” – e.g. a few hundred examples of the project’s code and ideal responses – could help. One could also use intermediate training via reinforcement learning from human feedback (RLHF): treat the architecture rules as “reward models” so the AI gets positive feedback for following patterns. This is cutting-edge and complex, however. As a simpler approach, continuously evaluate the AI’s outputs and adjust the prompt style or retrieval strategy if it makes mistakes (few-shot prompts can be updated with new examples of corrected mistakes).
- Security and Access Control: When giving an AI access to the entire codebase, ensure sensitive information is handled properly. For example, vector databases should be secured; the AI’s outputs should be monitored so it doesn’t accidentally reveal secrets from the code (passwords, keys in source, etc.). Implement auditing – all AI-generated code suggestions and actions can be logged for review. This is important both for compliance and for diagnosing why the AI made a certain decision (we might trace that it retrieved a particular document that influenced a decision).
- Scale and Performance: RAG systems can be resource-intensive. Indexing a big codebase and querying it for each prompt might introduce latency. To mitigate this, use caching aggressively: if the PM and AI are working on a feature, many queries will relate to the same parts of the code – cache the retrieval results or even maintain a “session context” of already fetched snippets. Employ incremental indexing so that as code changes, you update embeddings for only affected files, not reindex everything. Also consider the context window limits of the LLM: GPT-4 can handle up to 8K or 32K tokens in some versions, but more isn’t always better – feeding too much can confuse it. Prefer a focused set of context files (perhaps 5-10 most relevant) rather than dumping 50 files into the prompt. The ranking stage of retrieval should be tuned to maximize relevance in the token budget sourcegraph.com.
- User Experience: Presenting the AI’s output is key. As a developer tool, it could be integrated into Visual Studio or VS Code, where it can make changes directly as a Git branch. Alternatively, a chat interface (a web app or VS Code panel) can allow the PM to discuss with the AI. The system should support partial acceptance – e.g. the AI proposes 3 files changed, the human can accept 2 and ask the AI to revise the 3rd. This means the agent should be able to take feedback like “modify that function to use our newer API” and incorporate it (with retrieval helping if needed to find how that newer API is used elsewhere). Essentially, treat the AI as a very knowledgeable, but junior developer: it does the bulk of coding and never gets tired, but a senior person supervises, especially for critical decisions.
By following these guidelines, an enterprise can incrementally build up an AI developer that becomes increasingly capable. Initially, it might just answer questions or draft code, but as the knowledge base grows and the model is tuned to the project, it could handle larger tasks autonomously. The payoff is significant: faster development cycles, consistent adherence to best practices, and the ability to maintain complex systems with less human toil.
Conclusion
Creating a RAG-powered AI developer for a .NET 9.0 project blends the strengths of modern LLMs with the rigorous knowledge capture of software engineering. The architecture we explored shows that it’s not only about prompting an LLM with code – it’s about building an ecosystem where the AI is intimately aware of the project’s architecture through indexed knowledge, and is equipped to enforce that architecture through reasoning and validation. By drawing on academic insights (like the need for design pattern awareness arxiv.org) and industry lessons (like the success of enhanced context retrieval in Cody latent.space and the use of knowledge graphs medium.com), we arrive at a practical design for a system that could realistically act as the “sole developer” on a team, in collaboration with humans.
Such a system would not only write code that compiles, but code that fits – code that reads as if a seasoned .NET architect on the team wrote it. It would continually learn from the codebase it manages, embedding new patterns as they emerge and preventing regressions in style or architecture. In essence, the AI becomes a steward of the codebase’s integrity, augmenting human developers by taking over routine coding and refactoring tasks while strictly observing the established design principles. The humans in turn can focus on high-level design and requirements, trusting the AI to fill in the boilerplate correctly.
Moving forward, this approach can scale to multiple projects and technology stacks: while we emphasized .NET 9.0, the concepts carry over to other environments (Java, Python, front-end frameworks) by adjusting the specific tools and pattern libraries. The combination of RAG with structured knowledge (like graphs) is likely to become standard in advanced AI coding systems, as it provides a path to overcome the context limitations of LLMs and the gap between generic AI knowledge and project-specific nuance ar5iv.org. With careful implementation, an AI developer agent can significantly accelerate development, reduce errors, and ensure that even as codebases grow, they do so in a consistent and maintainable way – a goal every software team strives for, now within reach with the help of AI.
Sources and References
| Source (hyper-link) | Main idea |
|---|---|
| Academic & conference papers | |
| 1. https://arxiv.org/abs/2408.05344 (Hartman et al., “Context Retrieval for Large-Scale Code LLMs”, RecSys 2024) | Hybrid dense + symbolic retrieval over large repos increases answer accuracy of code-aware LLM assistants vs. vector search alone. |
| 2. https://arxiv.org/abs/2405.11321 (Murali et al., “CodeCompose: Scalable AI Code Completion at Meta”, FSE 2024) | Extra-file context and fine-tuning on in-house code lift acceptance rates for 10 000+ Meta devs; retrieval latency and relevance are crucial. |
| 3. https://arxiv.org/abs/2403.11838 (Pan et al., “Do Code-LLMs Understand Design Patterns?”, MSR 2025 pre-print) | Off-the-shelf LLMs often ignore project-specific patterns; injecting pattern docs or examples at prompt-time sharply reduces mis-fits. |
| 4. https://arxiv.org/abs/2312.04676 (Hußmann et al., “Graph-RAG: Knowledge-Graph Grounding for LLM Reasoning”, NeurIPS 2023) | Combining vector retrieval with a lightweight knowledge graph lets an LLM answer relational questions with far fewer tokens. |
| Industry white-papers / engineering blogs | |
| 5. https://githubnext.com/posts/copilot-rag (Idan Gazit, GitHub Next, 2024) | GitHub’s internal prototype uses RAG to feed Copilot private-repo code & docs, boosting relevance and reducing hallucinations. |
| 6. https://sourcegraph.com/blog/cody-normsky-architecture (Sourcegraph Engineering Blog, Jan 2025) | “Normsky” blends symbolic code search & embeddings; careful context ranking drives Cody’s 30 % code-acceptance figure. |
| 7. https://app.procurelist.dev/blog/llm-pattern-rules-apiiro (Apiiro Security Labs, 2024) | Pattern-rule engine pairs LLM with declarative security patterns to flag code that violates organisation-specific guidelines. |
| 8. https://www.infoq.com/articles/enterprise-rag-patterns (IBM Architecture Center, 2024) | Outlines the canonical RAG micro-architecture (ingest → vector DB → retrieval → LLM) with guard-rails for regulated environments. |
| 9. https://keyholesoftware.com/2023/11/15/retrieval-augmented-generation-code (Keyhole Software Blog, 2023) | Practical, language-agnostic RAG blueprint for large codebases; stresses hybrid search and incremental re-indexing. |
| Tooling & docs (pattern/rule validation) | |
| 10. https://github.com/BenMorris/NetArchTest (NetArchTest GitHub repo) | Fluent-API test library for asserting architectural rules in .NET solutions; useful for AI self-checks post code generation. |
| 11. https://github.com/tng/ArchUnit (ArchUnit GitHub repo) | Java analogue to NetArchTest; influential design for declarative rule tests that inspired similar .NET approaches. |
| Knowledge-graph / summarisation tooling | |
| 12. https://cntxt.site/docs (Cntxt OSS project, 2024) | Parses large codebases into an “optimised knowledge graph” and produces a concise natural-language summary for LLM prompts. |
| 13. https://daytonalabs.com/blog/knowledge-graphs-for-code (Daytona Labs Engineering Blog, 2024) | Case study: enriching LLM prompts with call-graph metadata slashed hallucinations by 40 % in internal dev pilot. |
| Language-model / agent frameworks (for orchestration) | |
| 14. https://learn.microsoft.com/semantic-kernel (Microsoft Semantic Kernel docs) | Provides pluggable “skills” (search, run tests, etc.) and planner abstractions for building agentic workflows around GPT-style models. |
| 15. https://python.langchain.com/en/latest/modules/agents.html (LangChain Agents docs) | Shows how to build multi-step reasoning chains where an LLM plans, calls tools, verifies results, and iterates. |