Customizing an LLM for Enterprise Software Engineering

[2605.16517] Customizing an LLM for Enterprise Software Engineering

-->

Computer Science > Software Engineering

arXiv:2605.16517 (cs)

[Submitted on 15 May 2026]

Title:Customizing an LLM for Enterprise Software Engineering

Authors:Aditya Kini, Satish Chandra, Milad Hashemi, Saksham Thakur, Aditya Pandey, Vincent Nguyen, Marc Brockschmidt, Franjo Ivančić, Danny Tarlow, Parthasarathy Ranganathan, Petros Maniatis, Ahmed Omran, Zaheer Abbas, Anita Gergely, Martin Sevenich, Gufeng Zhang, Amy Hua, Alexander Frömmgen Ranganathan View a PDF of the paper titled Customizing an LLM for Enterprise Software Engineering, by Aditya Kini and 17 other authors

View PDF HTML (experimental)

Abstract:Enterprise software development is a continuous evolutionary process, characterized by incremental additions, architectural revisions, production deployments and rigorous maintenance. These activities generate valuable data that modern LLMs could be finetuned on, to unlock additional tool possibilities for enterprise software engineering. While frontier LLMs are already very capable, this form of customization offers a compelling path for enterprise-specific optimization.

We introduce Gemini for Google (GfG)}, an adaptation of Gemini specialized for Google's internal software engineering ecosystem. This paper details the model's end-to-end development, from curating a trillion-token proprietary dataset to implementing a mid-training strategy that mitigates catastrophic forgetting. In a large-scale blind A/B study across 29,000 developers, Gemini for Google significantly outperformed baselines: reducing the mean number of iterations per turn by 23\%, and increasing code survival rates by about 17%. Beyond metrics, we provide a comprehensive blueprint for enterprise model adaptation, covering: (1)The extraction of high-value signals from software engineering data, (2)Data preparation strategies, (3)Full-stack model tuning (continued pre-training and post-training), and (4)The deployment of downstream applications. We believe this methodology offers a replicable path for other organizations to unlock the full potential of their internal engineering data.

Comments: 11 pages, 8 figures. To appear in ASE 2026 (Industry Track)

Subjects:

Software Engineering (cs.SE)

Cite as: arXiv:2605.16517 [cs.SE]

(or arXiv:2605.16517v1 [cs.SE] for this version)

https://doi.org/10.48550/arXiv.2605.16517

Focus to learn more

arXiv-issued DOI via DataCite

Submission history From: Saksham Thakur [view email] [v1] Fri, 15 May 2026 18:11:55 UTC (1,437 KB)

Full-text links: Access Paper:

View a PDF of the paper titled Customizing an LLM for Enterprise Software Engineering, by Aditya Kini and 17 other authors View PDF HTML (experimental) TeX Source

view license

Current browse context:

cs.SE

next >

new recent | 2026-05

Change to browse by:

References & Citations

NASA ADS Google Scholar

Semantic Scholar

export BibTeX citation Loading...

BibTeX formatted citation

Data provided by:

Bookmark

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Customizing an LLM for Enterprise Software Engineering

Related Articles

Elevated error rates on requests to multiple models

Donald Trump and sons to be 'forever' exempt from tax audits

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self- Play

Old Reddit Is Down

The ultimate female fantasy – A feminist critique of Beauty and the Beast