Skip to content

Dataset y resultados

Mindset & Code edited this page Aug 18, 2026 · 3 revisions

Dataset y resultados

🇬🇧 English first · 🇪🇸 Español más abajo.

Structure of sales_data.csv

10.000 rows, 9 columns, all of calendar year 2024 (first date 2024-01-01, last 2024-12-30).

Column Type How it is produced
SaleID int Sequential, 10000 … 19999
SaleDate date A random day inside 2024 — start + timedelta(randint(0, 364))
Region text North America 35% · Europe 30% · Asia 20% · LATAM 15%
ProductCategory text Software 50% · Hardware 30% · Service 20%
Revenue float Uniform between 100 and 5000
Cost float Uniform between 50 and 2000
SalesPersonID int 1 … 20
CustomerType text SMB 40% · Enterprise 30% · Individual 30%
Profit float Revenue − Cost, computed after the fact

The one thing to understand before reading any result

Revenue and Cost are drawn independently. Cost is not a function of revenue and carries no margin assumption. Three consequences follow, and they are visible in the data:

  1. 1.910 of the 10.000 sales close at a loss — whenever the cost draw lands above the revenue draw. That is 19% of the book.
  2. The global margin, 59,9%, is an artefact of the two ranges, not a business result. Mean revenue is 2.550,25 and mean cost is roughly 1.022.
  3. No segment carries a signal. Region, category and customer type differ from one another only by sampling noise, because nothing in the generator makes any of them more profitable.

Same for time: dates are uniform across the year, so there is no seasonality to find. A quarterly breakdown of this dataset returns four near-identical numbers.

This is a SQL exercise, not a simulation of a real business — and it is more useful to say so than to read meaning into the noise.

Real aggregates

Computed over the committed CSV:

Rows 10.000
Period 2024-01-01 → 2024-12-30
Total revenue 25.502.469,17
Total cost 10.224.911,27
Total profit 15.277.557,90
Global margin 59,91%
Average ticket 2.550,25
Average profit per sale 1.527,76
Sales at a loss 1.910
Salespeople 20

By region

Region Sales Revenue Profit Margin
North America 3.555 9.176.885,92 5.569.443,04 60,69%
Europe 3.048 7.728.074,08 4.596.612,65 59,48%
Asia 1.949 4.848.981,93 2.838.809,07 58,54%
LATAM 1.448 3.748.527,24 2.272.693,14 60,63%

The revenue ranking follows the sampling weights exactly. The margin spread — 58,54% to 60,69% — is two points of noise, not a finding.

By product category

Category Sales Avg revenue Avg cost Avg profit
Software 4.936 2.577,78 1.034,36 1.543,42
Hardware 3.031 2.545,49 1.004,51 1.540,99
Service 2.033 2.490,48 1.020,48 1.470,00

By customer type

Type Sales Revenue
SMB 3.946 10.134.770,65
Enterprise 3.057 7.839.138,01
Individual 2.997 7.528.560,51

A query that returns nothing, and why that is correct

Query 5 of sales_analysis.sql looks for categories whose average profit is under 1.000:

GROUP BY ProductCategory
HAVING AVG(Profit) < 1000;

With this dataset it returns an empty result set: the three categories average between 1.470 and 1.543. The query is written correctly; the threshold simply does not match data where cost has no relationship to revenue. Raising it above 1.550 would return all three, which is equally uninformative — the honest fix is to generate cost as a share of revenue, so that margin becomes something the data can actually differ on.


🇪🇸 Español

Estructura de sales_data.csv

10.000 filas, 9 columnas, todas del año natural 2024 (primera fecha 2024-01-01, última 2024-12-30).

Columna Tipo Cómo se produce
SaleID entero Secuencial, 10000 … 19999
SaleDate fecha Un día al azar dentro de 2024
Region texto North America 35% · Europe 30% · Asia 20% · LATAM 15%
ProductCategory texto Software 50% · Hardware 30% · Service 20%
Revenue decimal Uniforme entre 100 y 5000
Cost decimal Uniforme entre 50 y 2000
SalesPersonID entero 1 … 20
CustomerType texto SMB 40% · Enterprise 30% · Individual 30%
Profit decimal Revenue − Cost, calculado después

Lo único que hay que entender antes de leer cualquier resultado

Revenue y Cost se sortean de forma independiente. El coste no es función del ingreso y no lleva dentro ninguna hipótesis de margen. De ahí se siguen tres consecuencias, y se ven en los datos:

  1. 1.910 de las 10.000 ventas cierran en pérdidas, siempre que el sorteo del coste cae por encima del sorteo del ingreso. Es el 19% de la cartera.
  2. El margen global, 59,91%, es un artefacto de los dos rangos, no un resultado de negocio. El ingreso medio es 2.550,25 y el coste medio ronda 1.022.
  3. Ningún segmento lleva señal. Región, categoría y tipo de cliente se diferencian entre sí solo por ruido de muestreo, porque nada en el generador hace que ninguno sea más rentable.

Lo mismo con el tiempo: las fechas son uniformes a lo largo del año, así que no hay estacionalidad que encontrar. Un desglose trimestral de este conjunto devuelve cuatro cifras casi idénticas.

Esto es un ejercicio de SQL, no una simulación de un negocio real — y es más útil decirlo que leer significados en el ruido.

Agregados reales

Calculados sobre el CSV versionado.

Filas 10.000
Periodo 2024-01-01 → 2024-12-30
Ingreso total 25.502.469,17
Coste total 10.224.911,27
Beneficio total 15.277.557,90
Margen global 59,91%
Ticket medio 2.550,25
Beneficio medio por venta 1.527,76
Ventas en pérdidas 1.910
Comerciales 20

Por región

Región Ventas Ingreso Beneficio Margen
North America 3.555 9.176.885,92 5.569.443,04 60,69%
Europe 3.048 7.728.074,08 4.596.612,65 59,48%
Asia 1.949 4.848.981,93 2.838.809,07 58,54%
LATAM 1.448 3.748.527,24 2.272.693,14 60,63%

El orden por ingreso reproduce exactamente los pesos del muestreo. La horquilla de márgenes —del 58,54% al 60,69%— son dos puntos de ruido, no un hallazgo.

Por categoría de producto

Categoría Ventas Ingreso medio Coste medio Beneficio medio
Software 4.936 2.577,78 1.034,36 1.543,42
Hardware 3.031 2.545,49 1.004,51 1.540,99
Service 2.033 2.490,48 1.020,48 1.470,00

Por tipo de cliente

Tipo Ventas Ingreso
SMB 3.946 10.134.770,65
Enterprise 3.057 7.839.138,01
Individual 2.997 7.528.560,51

Una consulta que no devuelve nada, y por qué eso está bien

La quinta consulta de sales_analysis.sql busca categorías cuyo beneficio medio baje de 1.000. Con este conjunto de datos devuelve un resultado vacío: las tres categorías promedian entre 1.470 y 1.543. La consulta está bien escrita; lo que pasa es que el umbral no encaja con unos datos donde el coste no guarda relación con el ingreso. Subirlo por encima de 1.550 devolvería las tres, que informa lo mismo: la corrección honesta es generar el coste como porcentaje del ingreso, para que el margen pase a ser algo en lo que los datos puedan diferir de verdad.