Repository navigation
Home
Describe an Eloquent model once, and let an LLM write realistic seed data for it. AI fills the fields that need "intelligence" (names, blurbs, categories); Faker and literals fill everything else.
- PHP 8.2+ (Laravel 13 itself needs PHP 8.3+)
- Laravel 12 or 13
- An OpenAI API key, or any OpenAI-compatible server (Ollama, LM Studio, llama.cpp)
composer require betanow/laravel-ai-seeder
php artisan vendor:publish --tag=ai-seeder-configOpenAI:
AI_SEEDER_DRIVER=openai
OPENAI_API_KEY=sk-...
AI_SEEDER_OPENAI_MODEL=gpt-4o-miniA local OpenAI-compatible server (Ollama shown):
AI_SEEDER_DRIVER=local
AI_SEEDER_LOCAL_ENDPOINT=http://localhost:11434/v1/chat/completions
AI_SEEDER_LOCAL_MODEL=mistralBoth drivers speak the OpenAI chat-completions format and take the same settings (see config/ai-seeder.php). Each
setting except requires_api_key can be set with AI_SEEDER_OPENAI_<NAME> or AI_SEEDER_LOCAL_<NAME>, for example
AI_SEEDER_LOCAL_BATCH_SIZE:
| Setting | openai |
local |
|---|---|---|
api_key |
AI_SEEDER_OPENAI_API_KEY, else OPENAI_API_KEY
|
AI_SEEDER_LOCAL_API_KEY (optional) |
endpoint |
https://api.openai.com/v1/chat/completions |
http://localhost:11434/v1/chat/completions |
model |
gpt-4o-mini |
mistral |
batch_size |
10 rows per request | 10 |
max_concurrency |
4 requests at a time | 4 |
json_mode |
true |
false |
timeout |
60 s | 120 s |
retries |
3 | 3 |
retry_delay |
500 ms | 500 ms |
requires_api_key |
true |
false |
AI_SEEDER_DEFAULT_COUNT (default 10) is the number of rows per model when no count is given.
Create a definition class, for example app/Seeding/ProductDefinition.php:
<?php
namespace App\Seeding;
use Faker\Generator;
use Illuminate\Support\Str;
use Vendor\AiSeeder\Ai;
use Vendor\AiSeeder\AiModelDefinition;
class ProductDefinition extends AiModelDefinition
{
public function context (): string
{
return 'an online shop selling outdoor gear';
}
public function fields (): array
{
return [
'name' => Ai::text('product name'),
'description' => Ai::text('one-sentence marketing blurb'),
'category' => Ai::enum(['tents', 'boots', 'packs']),
'price' => Ai::float('price in EUR', 5, 500),
'in_stock' => Ai::boolean(),
'released_at' => Ai::date('release date'),
// Not sent to the LLM: a closure gets the Faker instance and the row built so far.
'slug' => fn (Generator $faker, array $row) => Str::slug($row['name']) . '-' . $faker->unique()->numberBetween(1000, 99999),
// Literals are used as they are.
'currency' => 'EUR',
'note' => NULL,
];
}
}Each field is exactly one of: an Ai::* descriptor (generated by the LLM and validated), a closure
fn (Generator $faker, array $row) (run after the AI values are known, in declaration order; $row holds every AI
value plus the local fields declared before it), or a literal.
Available descriptors: Ai::text, Ai::integer, Ai::float, Ai::boolean, Ai::date (Y-m-d), Ai::enum.
A definition without any Ai field never calls the LLM.
Only context() and the Ai descriptors (type, hint, options, range) are sent to the provider; closures and literals
stay in your app.
Generate unique columns (a slug, a SKU) with a Faker unique() closure, as slug does above, not from an Ai::*
field. Batches are requested concurrently, so two batches can return the same value for a unique column; the insert
then fails and the whole model is rolled back. unique() is finite (it throws an OverflowException once its range is
used up) and only remembers the values of the current run, not rows already in the table.
Register it. The order is the seeding order, so list parents before dependents:
// config/ai-seeder.php
'models' => [
App\Models\Product::class => App\Seeding\ProductDefinition::class,
],The command and the AiSeeder facade only work for models registered here; any other class is rejected.
php artisan ai-seeder:seed Product --count=25 # one model (class name or short name)
php artisan ai-seeder:seed --count=10 # every registered model, in order
php artisan ai-seeder:seed Product --dry-run # print rows, insert nothing (the provider is still called)
php artisan ai-seeder:seed Product --driver=local # use another driver for this run--count must be a positive integer. Leading zeros are fine (007 is 7); 0, negative numbers, decimals, +5,
values with spaces, non-numeric text and numbers beyond PHP's integer range are rejected with an error before anything
is sent. There is no maximum: every batch_size rows (default 10) is one request to the provider, and with OpenAI each
request is paid, so choose counts deliberately. 1,000 rows is at least 100 requests; that is a floor, because corrective
requests (see "How it works") and HTTP retries add more.
From code or a DatabaseSeeder:
use App\Models\Product;
use Vendor\AiSeeder\Facades\AiSeeder;
AiSeeder::seed(Product::class, 25); // Collection of created models
AiSeeder::generate(Product::class, 5); // rows only, nothing written- The count is split by
batch_size; batches are sentmax_concurrencyat a time. - Each reply must be JSON (
{"rows": [...]}, a bare array, or a fenced block). Rows are validated against theAifields; invalid rows are dropped and the shortfall is requested once more. Still short meansGenerationFailed. - Rows are written with
forceCreate(casts, mutators, events and timestamps apply).
- The rows of one model are written in one transaction, on the model's own connection, and only after every batch has
been generated. If generation fails, or the insert fails (for example on a unique column), nothing of that model is
written; a database error during the insert is reported as
GenerationFailed. - With several models, models finished earlier stay written if a later model fails during generation or insert.
- Misconfigured targets are rejected before anything is generated or written: an unknown model, a definition class
that does not extend
AiModelDefinition, or (in real runs, not with--dry-run) a model class that is not Eloquent. A typo in a registered entry therefore fails the whole run up front.
creating / created observers fire inside the transaction. An observer with side effects (mail, queued jobs) should
set public $afterCommit = TRUE; (or implement ShouldHandleEventsAfterCommit), so it only runs once the rows are
committed.
HTTP 429, HTTP 5xx and connection errors are retried with exponential backoff: the wait starts at retry_delay
milliseconds (default 500) and doubles each time, up to retries retries (default 3). Other 4xx errors are not
retried. The API key never appears in error messages.
-
The [openai] AI driver needs an API key ...: setOPENAI_API_KEY(orAI_SEEDER_OPENAI_API_KEY). Nothing was sent. Thelocaldriver does not need a key. -
GenerationFailedsaying only some rows were usable "after one corrective retry": the model kept returning rows that were not valid JSON or did not match the field types. Try a stronger model, or lowerbatch_size. - A local server must expose an OpenAI-compatible
/v1/chat/completionsendpoint. For Ollama that ishttp://localhost:11434/v1/chat/completions; setAI_SEEDER_LOCAL_MODELto a model the server has pulled. - A local server that handles one request at a time (a single-slot llama.cpp, for example) can time out or fail when
max_concurrencyrequests arrive together: setAI_SEEDER_LOCAL_MAX_CONCURRENCY=1, and raiseAI_SEEDER_LOCAL_TIMEOUT(seconds) if single requests are slow. -
json_modeasks the provider for a JSON object (response_format). It defaults totrueforopenaiandfalseforlocal, because not every compatible server supports it; enable it withAI_SEEDER_LOCAL_JSON_MODE=trueif yours does. -
HTTP 4xxfrom the provider (other than 429) is not retried; read the provider's message in the error and check the key, model name and endpoint.
vendor/bin/phpunit # unit + feature + smoke; no network, no API key
vendor/bin/phpunit --group smoke # only the smoke tests
vendor/bin/pint --test # code styleThe smoke tests make real HTTP calls to a local fake OpenAI server. They start PHP's built-in server with php -S
and use pgrep to clean up its workers, so they need macOS or Linux.
One opt-in test calls the real OpenAI API and makes a paid request. phpunit.xml excludes its live group from the
default run, so it only runs when you select the group, and even then it is skipped unless both variables are set; an
exported key alone never triggers it:
AI_SEEDER_LIVE_TEST=1 OPENAI_API_KEY=sk-... vendor/bin/phpunit --group liveA runnable demo (an outdoor-gear Product model) lives in workbench/ and runs through Testbench, against the fake
provider used by the smoke tests. Run it from the package root after composer install.
The demo database is vendor/orchestra/testbench-core/laravel/database/demo.sqlite; the workbench provider creates it
on demand. Always start with migrate:fresh: it rebuilds the products table and also Testbench's own users, cache
and jobs tables in that database. Wiping vendor/ or running composer update deletes the demo database, and
migrate:fresh rebuilds it.
vendor/bin/testbench migrate:freshIn a second terminal, start the fake provider and leave it running (stop it with Ctrl+C when done):
php -S 127.0.0.1:8089 tests/Support/fake-openai-server.phpThen seed and look at the result:
AI_SEEDER_DRIVER=local AI_SEEDER_LOCAL_ENDPOINT=http://127.0.0.1:8089/v1/chat/completions \
vendor/bin/testbench ai-seeder:seed Product --count=25
sqlite3 -header -column vendor/orchestra/testbench-core/laravel/database/demo.sqlite \
"select id, name, category, price, in_stock, released_at, slug from products limit 5"The sqlite3 command needs the SQLite command-line tool to be installed. Dates show as YYYY-MM-DD 00:00:00 because
the Product model casts released_at with date.
Relationships between models are not resolved for you (a closure may look up existing ids itself); there is no queue or streaming support; only OpenAI-style chat-completions servers are supported.