-
Notifications
You must be signed in to change notification settings - Fork 0
Dataset y resultados
🇬🇧 English first · 🇪🇸 Español más abajo.
10.000 rows, 9 columns, all of calendar year 2024 (first date 2024-01-01, last 2024-12-30).
| Column | Type | How it is produced |
|---|---|---|
SaleID |
int | Sequential, 10000 … 19999 |
SaleDate |
date | A random day inside 2024 — start + timedelta(randint(0, 364))
|
Region |
text | North America 35% · Europe 30% · Asia 20% · LATAM 15% |
ProductCategory |
text | Software 50% · Hardware 30% · Service 20% |
Revenue |
float | Uniform between 100 and 5000 |
Cost |
float | Uniform between 50 and 2000 |
SalesPersonID |
int | 1 … 20 |
CustomerType |
text | SMB 40% · Enterprise 30% · Individual 30% |
Profit |
float |
Revenue − Cost, computed after the fact |
Revenue and Cost are drawn independently. Cost is not a function of revenue
and carries no margin assumption. Three consequences follow, and they are visible in
the data:
- 1.910 of the 10.000 sales close at a loss — whenever the cost draw lands above the revenue draw. That is 19% of the book.
- The global margin, 59,9%, is an artefact of the two ranges, not a business result. Mean revenue is 2.550,25 and mean cost is roughly 1.022.
- No segment carries a signal. Region, category and customer type differ from one another only by sampling noise, because nothing in the generator makes any of them more profitable.
Same for time: dates are uniform across the year, so there is no seasonality to find. A quarterly breakdown of this dataset returns four near-identical numbers.
This is a SQL exercise, not a simulation of a real business — and it is more useful to say so than to read meaning into the noise.
Computed over the committed CSV:
| Rows | 10.000 |
| Period | 2024-01-01 → 2024-12-30 |
| Total revenue | 25.502.469,17 |
| Total cost | 10.224.911,27 |
| Total profit | 15.277.557,90 |
| Global margin | 59,91% |
| Average ticket | 2.550,25 |
| Average profit per sale | 1.527,76 |
| Sales at a loss | 1.910 |
| Salespeople | 20 |
| Region | Sales | Revenue | Profit | Margin |
|---|---|---|---|---|
| North America | 3.555 | 9.176.885,92 | 5.569.443,04 | 60,69% |
| Europe | 3.048 | 7.728.074,08 | 4.596.612,65 | 59,48% |
| Asia | 1.949 | 4.848.981,93 | 2.838.809,07 | 58,54% |
| LATAM | 1.448 | 3.748.527,24 | 2.272.693,14 | 60,63% |
The revenue ranking follows the sampling weights exactly. The margin spread — 58,54% to 60,69% — is two points of noise, not a finding.
| Category | Sales | Avg revenue | Avg cost | Avg profit |
|---|---|---|---|---|
| Software | 4.936 | 2.577,78 | 1.034,36 | 1.543,42 |
| Hardware | 3.031 | 2.545,49 | 1.004,51 | 1.540,99 |
| Service | 2.033 | 2.490,48 | 1.020,48 | 1.470,00 |
| Type | Sales | Revenue |
|---|---|---|
| SMB | 3.946 | 10.134.770,65 |
| Enterprise | 3.057 | 7.839.138,01 |
| Individual | 2.997 | 7.528.560,51 |
Query 5 of sales_analysis.sql looks for categories whose average profit is under
1.000:
GROUP BY ProductCategory
HAVING AVG(Profit) < 1000;With this dataset it returns an empty result set: the three categories average between 1.470 and 1.543. The query is written correctly; the threshold simply does not match data where cost has no relationship to revenue. Raising it above 1.550 would return all three, which is equally uninformative — the honest fix is to generate cost as a share of revenue, so that margin becomes something the data can actually differ on.
10.000 filas, 9 columnas, todas del año natural 2024 (primera fecha 2024-01-01, última 2024-12-30).
| Columna | Tipo | Cómo se produce |
|---|---|---|
SaleID |
entero | Secuencial, 10000 … 19999 |
SaleDate |
fecha | Un día al azar dentro de 2024 |
Region |
texto | North America 35% · Europe 30% · Asia 20% · LATAM 15% |
ProductCategory |
texto | Software 50% · Hardware 30% · Service 20% |
Revenue |
decimal | Uniforme entre 100 y 5000 |
Cost |
decimal | Uniforme entre 50 y 2000 |
SalesPersonID |
entero | 1 … 20 |
CustomerType |
texto | SMB 40% · Enterprise 30% · Individual 30% |
Profit |
decimal |
Revenue − Cost, calculado después |
Revenue y Cost se sortean de forma independiente. El coste no es función del
ingreso y no lleva dentro ninguna hipótesis de margen. De ahí se siguen tres
consecuencias, y se ven en los datos:
- 1.910 de las 10.000 ventas cierran en pérdidas, siempre que el sorteo del coste cae por encima del sorteo del ingreso. Es el 19% de la cartera.
- El margen global, 59,91%, es un artefacto de los dos rangos, no un resultado de negocio. El ingreso medio es 2.550,25 y el coste medio ronda 1.022.
- Ningún segmento lleva señal. Región, categoría y tipo de cliente se diferencian entre sí solo por ruido de muestreo, porque nada en el generador hace que ninguno sea más rentable.
Lo mismo con el tiempo: las fechas son uniformes a lo largo del año, así que no hay estacionalidad que encontrar. Un desglose trimestral de este conjunto devuelve cuatro cifras casi idénticas.
Esto es un ejercicio de SQL, no una simulación de un negocio real — y es más útil decirlo que leer significados en el ruido.
Calculados sobre el CSV versionado.
| Filas | 10.000 |
| Periodo | 2024-01-01 → 2024-12-30 |
| Ingreso total | 25.502.469,17 |
| Coste total | 10.224.911,27 |
| Beneficio total | 15.277.557,90 |
| Margen global | 59,91% |
| Ticket medio | 2.550,25 |
| Beneficio medio por venta | 1.527,76 |
| Ventas en pérdidas | 1.910 |
| Comerciales | 20 |
| Región | Ventas | Ingreso | Beneficio | Margen |
|---|---|---|---|---|
| North America | 3.555 | 9.176.885,92 | 5.569.443,04 | 60,69% |
| Europe | 3.048 | 7.728.074,08 | 4.596.612,65 | 59,48% |
| Asia | 1.949 | 4.848.981,93 | 2.838.809,07 | 58,54% |
| LATAM | 1.448 | 3.748.527,24 | 2.272.693,14 | 60,63% |
El orden por ingreso reproduce exactamente los pesos del muestreo. La horquilla de márgenes —del 58,54% al 60,69%— son dos puntos de ruido, no un hallazgo.
| Categoría | Ventas | Ingreso medio | Coste medio | Beneficio medio |
|---|---|---|---|---|
| Software | 4.936 | 2.577,78 | 1.034,36 | 1.543,42 |
| Hardware | 3.031 | 2.545,49 | 1.004,51 | 1.540,99 |
| Service | 2.033 | 2.490,48 | 1.020,48 | 1.470,00 |
| Tipo | Ventas | Ingreso |
|---|---|---|
| SMB | 3.946 | 10.134.770,65 |
| Enterprise | 3.057 | 7.839.138,01 |
| Individual | 2.997 | 7.528.560,51 |
La quinta consulta de sales_analysis.sql busca categorías cuyo beneficio medio
baje de 1.000. Con este conjunto de datos devuelve un resultado vacío: las tres
categorías promedian entre 1.470 y 1.543. La consulta está bien escrita; lo que
pasa es que el umbral no encaja con unos datos donde el coste no guarda relación con
el ingreso. Subirlo por encima de 1.550 devolvería las tres, que informa lo mismo:
la corrección honesta es generar el coste como porcentaje del ingreso, para que el
margen pase a ser algo en lo que los datos puedan diferir de verdad.