Skip to content
Evan Grosset-Janin edited this page Oct 8, 2026 · 2 revisions

Laravel AI Seeder

Describe an Eloquent model once, and let an LLM write realistic seed data for it. AI fills the fields that need "intelligence" (names, blurbs, categories); Faker and literals fill everything else.

Requirements

  • PHP 8.2+ (Laravel 13 itself needs PHP 8.3+)
  • Laravel 12 or 13
  • An OpenAI API key, or any OpenAI-compatible server (Ollama, LM Studio, llama.cpp)

Installation

composer require betanow/laravel-ai-seeder
php artisan vendor:publish --tag=ai-seeder-config

Configure a driver

OpenAI:

AI_SEEDER_DRIVER=openai
OPENAI_API_KEY=sk-...
AI_SEEDER_OPENAI_MODEL=gpt-4o-mini

A local OpenAI-compatible server (Ollama shown):

AI_SEEDER_DRIVER=local
AI_SEEDER_LOCAL_ENDPOINT=http://localhost:11434/v1/chat/completions
AI_SEEDER_LOCAL_MODEL=mistral

Both drivers speak the OpenAI chat-completions format and take the same settings (see config/ai-seeder.php). Each setting except requires_api_key can be set with AI_SEEDER_OPENAI_<NAME> or AI_SEEDER_LOCAL_<NAME>, for example AI_SEEDER_LOCAL_BATCH_SIZE:

Setting openai local
api_key AI_SEEDER_OPENAI_API_KEY, else OPENAI_API_KEY AI_SEEDER_LOCAL_API_KEY (optional)
endpoint https://api.openai.com/v1/chat/completions http://localhost:11434/v1/chat/completions
model gpt-4o-mini mistral
batch_size 10 rows per request 10
max_concurrency 4 requests at a time 4
json_mode true false
timeout 60 s 120 s
retries 3 3
retry_delay 500 ms 500 ms
requires_api_key true false

AI_SEEDER_DEFAULT_COUNT (default 10) is the number of rows per model when no count is given.

Describe a model

Create a definition class, for example app/Seeding/ProductDefinition.php:

<?php

namespace App\Seeding;

use Faker\Generator;
use Illuminate\Support\Str;
use Vendor\AiSeeder\Ai;
use Vendor\AiSeeder\AiModelDefinition;

class ProductDefinition extends AiModelDefinition
{
    public function context (): string
    {
        return 'an online shop selling outdoor gear';
    }

    public function fields (): array
    {
        return [
            'name' => Ai::text('product name'),
            'description' => Ai::text('one-sentence marketing blurb'),
            'category' => Ai::enum(['tents', 'boots', 'packs']),
            'price' => Ai::float('price in EUR', 5, 500),
            'in_stock' => Ai::boolean(),
            'released_at' => Ai::date('release date'),

            // Not sent to the LLM: a closure gets the Faker instance and the row built so far.
            'slug' => fn (Generator $faker, array $row) => Str::slug($row['name']) . '-' . $faker->unique()->numberBetween(1000, 99999),

            // Literals are used as they are.
            'currency' => 'EUR',
            'note' => NULL,
        ];
    }
}

Each field is exactly one of: an Ai::* descriptor (generated by the LLM and validated), a closure fn (Generator $faker, array $row) (run after the AI values are known, in declaration order; $row holds every AI value plus the local fields declared before it), or a literal. Available descriptors: Ai::text, Ai::integer, Ai::float, Ai::boolean, Ai::date (Y-m-d), Ai::enum. A definition without any Ai field never calls the LLM.

Only context() and the Ai descriptors (type, hint, options, range) are sent to the provider; closures and literals stay in your app.

Generate unique columns (a slug, a SKU) with a Faker unique() closure, as slug does above, not from an Ai::* field. Batches are requested concurrently, so two batches can return the same value for a unique column; the insert then fails and the whole model is rolled back. unique() is finite (it throws an OverflowException once its range is used up) and only remembers the values of the current run, not rows already in the table.

Register it. The order is the seeding order, so list parents before dependents:

// config/ai-seeder.php
'models' => [
    App\Models\Product::class => App\Seeding\ProductDefinition::class,
],

The command and the AiSeeder facade only work for models registered here; any other class is rejected.

Seed

php artisan ai-seeder:seed Product --count=25     # one model (class name or short name)
php artisan ai-seeder:seed --count=10             # every registered model, in order
php artisan ai-seeder:seed Product --dry-run      # print rows, insert nothing (the provider is still called)
php artisan ai-seeder:seed Product --driver=local # use another driver for this run

--count must be a positive integer. Leading zeros are fine (007 is 7); 0, negative numbers, decimals, +5, values with spaces, non-numeric text and numbers beyond PHP's integer range are rejected with an error before anything is sent. There is no maximum: every batch_size rows (default 10) is one request to the provider, and with OpenAI each request is paid, so choose counts deliberately. 1,000 rows is at least 100 requests; that is a floor, because corrective requests (see "How it works") and HTTP retries add more.

From code or a DatabaseSeeder:

use App\Models\Product;
use Vendor\AiSeeder\Facades\AiSeeder;

AiSeeder::seed(Product::class, 25);      // Collection of created models
AiSeeder::generate(Product::class, 5);   // rows only, nothing written

How it works

  1. The count is split by batch_size; batches are sent max_concurrency at a time.
  2. Each reply must be JSON ({"rows": [...]}, a bare array, or a fenced block). Rows are validated against the Ai fields; invalid rows are dropped and the shortfall is requested once more. Still short means GenerationFailed.
  3. Rows are written with forceCreate (casts, mutators, events and timestamps apply).

Atomicity

  • The rows of one model are written in one transaction, on the model's own connection, and only after every batch has been generated. If generation fails, or the insert fails (for example on a unique column), nothing of that model is written; a database error during the insert is reported as GenerationFailed.
  • With several models, models finished earlier stay written if a later model fails during generation or insert.
  • Misconfigured targets are rejected before anything is generated or written: an unknown model, a definition class that does not extend AiModelDefinition, or (in real runs, not with --dry-run) a model class that is not Eloquent. A typo in a registered entry therefore fails the whole run up front.

Model events

creating / created observers fire inside the transaction. An observer with side effects (mail, queued jobs) should set public $afterCommit = TRUE; (or implement ShouldHandleEventsAfterCommit), so it only runs once the rows are committed.

Retries

HTTP 429, HTTP 5xx and connection errors are retried with exponential backoff: the wait starts at retry_delay milliseconds (default 500) and doubles each time, up to retries retries (default 3). Other 4xx errors are not retried. The API key never appears in error messages.

Troubleshooting

  • The [openai] AI driver needs an API key ...: set OPENAI_API_KEY (or AI_SEEDER_OPENAI_API_KEY). Nothing was sent. The local driver does not need a key.
  • GenerationFailed saying only some rows were usable "after one corrective retry": the model kept returning rows that were not valid JSON or did not match the field types. Try a stronger model, or lower batch_size.
  • A local server must expose an OpenAI-compatible /v1/chat/completions endpoint. For Ollama that is http://localhost:11434/v1/chat/completions; set AI_SEEDER_LOCAL_MODEL to a model the server has pulled.
  • A local server that handles one request at a time (a single-slot llama.cpp, for example) can time out or fail when max_concurrency requests arrive together: set AI_SEEDER_LOCAL_MAX_CONCURRENCY=1, and raise AI_SEEDER_LOCAL_TIMEOUT (seconds) if single requests are slow.
  • json_mode asks the provider for a JSON object (response_format). It defaults to true for openai and false for local, because not every compatible server supports it; enable it with AI_SEEDER_LOCAL_JSON_MODE=true if yours does.
  • HTTP 4xx from the provider (other than 429) is not retried; read the provider's message in the error and check the key, model name and endpoint.

Testing

vendor/bin/phpunit                  # unit + feature + smoke; no network, no API key
vendor/bin/phpunit --group smoke    # only the smoke tests
vendor/bin/pint --test              # code style

The smoke tests make real HTTP calls to a local fake OpenAI server. They start PHP's built-in server with php -S and use pgrep to clean up its workers, so they need macOS or Linux.

One opt-in test calls the real OpenAI API and makes a paid request. phpunit.xml excludes its live group from the default run, so it only runs when you select the group, and even then it is skipped unless both variables are set; an exported key alone never triggers it:

AI_SEEDER_LIVE_TEST=1 OPENAI_API_KEY=sk-... vendor/bin/phpunit --group live

Demo

A runnable demo (an outdoor-gear Product model) lives in workbench/ and runs through Testbench, against the fake provider used by the smoke tests. Run it from the package root after composer install.

The demo database is vendor/orchestra/testbench-core/laravel/database/demo.sqlite; the workbench provider creates it on demand. Always start with migrate:fresh: it rebuilds the products table and also Testbench's own users, cache and jobs tables in that database. Wiping vendor/ or running composer update deletes the demo database, and migrate:fresh rebuilds it.

vendor/bin/testbench migrate:fresh

In a second terminal, start the fake provider and leave it running (stop it with Ctrl+C when done):

php -S 127.0.0.1:8089 tests/Support/fake-openai-server.php

Then seed and look at the result:

AI_SEEDER_DRIVER=local AI_SEEDER_LOCAL_ENDPOINT=http://127.0.0.1:8089/v1/chat/completions \
    vendor/bin/testbench ai-seeder:seed Product --count=25

sqlite3 -header -column vendor/orchestra/testbench-core/laravel/database/demo.sqlite \
    "select id, name, category, price, in_stock, released_at, slug from products limit 5"

The sqlite3 command needs the SQLite command-line tool to be installed. Dates show as YYYY-MM-DD 00:00:00 because the Product model casts released_at with date.

Limitations

Relationships between models are not resolved for you (a closure may look up existing ids itself); there is no queue or streaming support; only OpenAI-style chat-completions servers are supported.