AI for Research and Marketing

NVIDIA Released 148,000 Records of Synthetic Salvadoran Personas — Here Is What We Found at Liorant

Research article Updated 2026 12 min read By Ricardo Mendoza Castro

NVIDIA's Nemotron-Personas-El-Salvador dataset places 148,000 records of synthetic Salvadoran personas directly in the hands of any analyst, marketer, or researcher willing to look. Each record contains seven distinct persona narratives — professional life, sports habits, artistic interests, travel behaviour, culinary preferences, family dynamics, and a general profile — adding up to roughly one million individual persona texts across the full dataset.

The dataset at a glance

148,000
records published by NVIDIA
× 7 persona texts per record
≈ 1,000,000
persona texts in total

Personas by department

27%13%9%8%8%6%5%5%4%4%3%3%3%2%
Circle size & shade = persona share

San Salvador alone holds 26.8% of the synthetic population; all 14 departments are represented.

25 fields, three tiers

7 Persona narrative fields
6 Contextual text fields
11 Demographic fields + UUID
25 columns · 0% missing (Liorant analysis) · single training split · Licence: CC-BY 4.0

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

NVIDIA, in collaboration with Widelabs, designed the dataset for Sovereign AI research — to support LLM training, improve synthetic data diversity, and reduce model collapse. This article explores how research teams can use it as an audience hypothesis-generation layer; that is Liorant's own applied interpretation, not NVIDIA's stated intent.

Liorant ran this dataset through a full analysis pipeline: exploratory statistics across all 148,000 rows, AI-generated embeddings of the narrative fields, an unsupervised segmentation of a 20,000-persona working sample, and a conversational marketing-research demo powered by a small language model. This article presents what the analysis surfaced — findings, patterns, and possible audience hypotheses — and explains what marketing teams and executives can do with a workflow like this.

One framing note before we begin: everything produced from this dataset is exploratory by nature. Synthetic personas represent plausible demographic and narrative diversity, not real consumer behaviour.

Patterns emerging from a 20,000-row sample point toward hypotheses worth testing, not conclusions to act on — which is why Liorant believes this kind of analysis is most valuable at the beginning of an audience-research process, not the end.

What the Nemotron-Personas-El-Salvador dataset contains

The dataset is available at nvidia/Nemotron-Personas-El-Salvador on Hugging Face, distributed in Parquet format as a single training split. It contains 148,000 rows and 25 columns.

Structured demographic fields provide the factual scaffolding for each persona: sex, age, languages spoken, marital status, household type, education level, occupation, urban or rural classification, municipality, department, and country. The occupation field covers 119 CIIU Rev.4 economic-activity classes across eight census employment categories.

Text fields give each persona its depth. NVIDIA's documentation identifies seven core persona narrative fields per record, which together produce approximately one million individual persona texts across the full dataset. The schema also includes six additional contextual text fields. Based on Liorant's analysis, all 13 text columns are unique across all 148,000 rows, with an average combined narrative length of approximately 6,817 characters and 1,165 words per record.

Every field across all 25 columns shows zero missing values. The dataset is structurally complete — every persona can be used without imputation, filtering, or data cleaning.

Diagram

Dataset architecture: 25 fields in three tiers

7 Core persona narrative fields
  • personaGeneral profile
  • professional_personaCareer & work
  • sports_personaPhysical activity
  • arts_personaCultural interests
  • travel_personaTravel behaviour
  • culinary_personaFood & eating
  • family_personaHousehold & family
≈ 1M unique texts across the full dataset.
6 Contextual text fields
  • cultural_background
  • skills_and_expertise
  • skills_and_expertise_list
  • hobbies_and_interests
  • hobbies_and_interests_list
  • career_goals_and_ambitions
Prose paragraphs + structured JSON list variants.
11 Structured demographic fields + UUID
  • uuid · sex · age
  • languages_spoken
  • marital_status · household_type
  • education_level · occupation
  • area (urban / rural)
  • municipality · department · country
Categorical and numeric values. Zero missing across all rows (Liorant analysis).
Total: 25 columns · 148,000 rows · single training split · Licence: CC-BY 4.0

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Exploring 148,000 personas without overloading your systems

Loading the full dataset directly into a standard analytics environment would require several gigabytes of memory — beyond what most shared or cloud notebook environments handle comfortably. Liorant's approach in the first phase was to query the dataset's files remotely using an in-memory SQL engine, without ever loading the complete dataset into a local environment.

This produced all descriptive statistics — row counts, schema inspection, missing-value counts, unique-value counts, and top categorical frequencies — through direct queries against the files as they exist on Hugging Face. The only data pulled into memory locally was a 20,000-row random sample, extracted for the downstream segmentation and AI-demo work.

The result: a complete demographic and structural overview of all 148,000 personas, produced without the infrastructure costs or memory requirements that normally accompany a dataset of this size.

Key structural finding (Liorant analysis). Every one of the 25 columns registers 0% missing values. All 13 text columns each contain 148,000 unique entries — every persona holds a fully distinct text in every field. From a data-quality standpoint, the dataset is unusually clean.

Workflow

The three-notebook analysis pipeline

Notebook 1 · Remote EDA
SQL queries run against Hugging Face Parquet files — the full dataset never leaves the cloud.
20,000-row sample + EDA stats
Notebook 2 · Segmentation
Multilingual AI model embeds persona texts into 384-dimensional vectors · k=4 clustering applied.
Embeddings + clustered sample
Notebook 3 · Marketing demo
Semantic search retrieves relevant personas · a small language model generates structured hypotheses.
Browser-based Q&A interface
Runs entirely on the Google Colab free tier · No cloud required: Optimized for local execution

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

What the data shows: key demographic and occupational patterns

The exploratory analysis of all 148,000 personas reveals several structural patterns. These are descriptive findings from the synthetic dataset — direct counts and distributions — not inferences about real Salvadoran consumers.

Age

The synthetic population spans ages 18 to 100. The mean age sits at 42.8 years (standard deviation 17.6); the median is 40. The interquartile range runs from 28 to 55 — half the population falls between those ages. The distribution is broader than a typical young-adult marketing sample and more representative of the full working-age-to-elderly spectrum.

Visual 1

Age distribution — Nemotron Personas El Salvador (n = 20,000 sample)

Median 40 · Mean 42.8183040557085100AGE

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Sex

A slight female majority — 79,632 female personas (53.81%) to 68,368 male (46.19%). This reflects design decisions in the dataset rather than census proportions, though it is broadly consistent with El Salvador's demographic structure.

Visual 1b

Sex distribution — full dataset

148,000 personas Female · 53.81%79,632 personas Male · 46.19%68,368 personas

Slight female majority reflects dataset design choices, not census data.

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Languages spoken

LanguageCount%
Spanish only134,65090.98%
Spanish and English12,7708.63%
Spanish and another language4530.31%
Salvadoran Sign Language (LESSA)890.06%
Spanish and indigenous language370.03%
Indigenous language only10.00%

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

For any team targeting El Salvador, this confirms that Spanish-only creative reaches the overwhelming majority of any realistic audience. Bilingual Spanish-English content is relevant for under 9% of the synthetic population, concentrated in groups with likely higher formal education and urban residence.

Visual 1c

Languages spoken — synthetic population

Spanish only9 in 10 personas speak only Spanish90.98%
Spanish + English8.63%
All others combined0.39%

Bilingual Spanish-English personas cluster in urban, higher-education profiles.

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Education level

LevelCount%
Bachillerato (high school)45,43630.70%
Primaria (primary)36,43724.62%
Secundaria (middle school)26,05717.61%
Universitario (university)17,76612.00%
Ninguno (no formal education)17,36311.73%
Técnico (technical)3,8642.61%
Posgrado (postgraduate)1,0770.73%

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

More than half the synthetic population holds at most a primary or middle-school education. University-level education appears in just 12% of personas. This has direct implications for how brands calibrate message complexity, financial-product design, digital onboarding, and trust-building across different audience groups.

Visual 2

Education level distribution — 148,000 synthetic personas

Bachillerato45,436 · 30.70%
Primaria36,437 · 24.62%
Secundaria26,057 · 17.61%
Universitario17,766 · 12.00%
Ninguno17,363 · 11.73%
Técnico3,864 · 2.61%
Posgrado1,077 · 0.73%

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Department — geographic distribution

San Salvador concentrates 26.76% of the synthetic population. La Libertad, Santa Ana, and Sonsonate add another 30.8%. All 14 departments appear, with the remaining ten collectively accounting for roughly 42% of personas. Any brand treating El Salvador as a single market centred on the capital is, by the dataset's own distribution, leaving the majority of the audience unaddressed.

Visual 3

Synthetic personas by department

San Salvador39,603 · 26.76%
La Libertad19,625 · 13.26%
Santa Ana13,868 · 9.37%
Sonsonate12,111 · 8.18%
San Miguel11,140 · 7.53%
Ahuachapán8,728 · 5.90%
Usulután8,094 · 5.47%
La Paz7,737 · 5.23%
Cuscatlán6,115 · 4.13%
La Unión5,337 · 3.61%
Chalatenango4,261 · 2.88%
Morazán4,005 · 2.71%
San Vicente3,902 · 2.64%
Cabañas3,474 · 2.35%

San Salvador concentrates the population; the remaining 13 departments fragment the rest.

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Top occupations

The dataset covers 119 CIIU Rev.4 economic-activity classes. The seven most frequent occupations in the full dataset, based on Liorant's analysis:

OccupationCount%
Food services and street food12,3598.35%
Grain and legume cultivation9,3246.30%
Non-specialised retail (food-led)9,1636.19%
Domestic household employment8,9966.08%
Building construction6,8624.64%
Bakery product manufacturing6,1374.15%
Garment manufacturing5,0683.42%

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

The economy represented skews toward informal commerce, agriculture, domestic work, and food production — consistent with the structure of El Salvador's real labour market, where the informal sector accounts for a substantial share of employment. Brands designing campaigns around formal, salaried workers will, by these distributions, be optimising for a minority profile.

We analyzed four audience patterns emerging from the segmentation analysis

The second phase moved from full-dataset descriptive statistics to a working sample of 20,000 randomly selected personas. The goal: test whether AI-driven segmentation of the narrative text fields produces coherent audience groupings that go beyond what structured demographic columns alone reveal.

The approach combined the seven official persona fields and the available contextual text fields into a single text representation, then used a multilingual AI model to convert each persona's combined narrative into a numerical vector — a mathematical representation of its meaning. Those vectors were grouped using an unsupervised algorithm, tested at multiple group sizes to identify the most internally coherent solution.

Workflow

How the audience segments were built

20,000-persona working sample
Combine 7 persona fields + contextual text into one block per record
Multilingual AI model encodes each block as a 384-number vector
Test k = 3–10 — measure internal cohesion
k = 4 selected — highest cohesion (0.119)
Profile each segment across demographics, geography, occupation
All steps represent Liorant's analysis · results are exploratory audience hypotheses, not statistically validated customer segments

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

A statistical measure of segment cohesion — how clearly separated and internally consistent each grouping is — peaked at four segments, scoring 0.1185 on a scale where higher is better. That score is modest in absolute terms, as expected for rich narrative text rather than purely numerical data, but it meaningfully outperforms every other group count tested (three through ten).

Visual 4

Segment cohesion score by number of audience groups

0.060.080.100.12Optimal: 4 segments (0.119)345678910NUMBER OF AUDIENCE GROUPS (k)

Higher scores indicate more internally consistent groupings.

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

The four segments below represent patterns that surfaced from the narrative analysis of the working sample. They are audience hypotheses grounded in the data, not verified customer profiles. Their value is in structuring the questions a team should take into real research — not in replacing that research.

Segment A — Informal service workers (28.8% · 5,757 personas)

The largest segment by volume and the most gender-specific in the dataset. 97.99% of personas in this segment are female. Occupations cluster around domestic household employment (831), street food and restaurant work (509), non-specialised retail (485), and garment manufacturing (328). Education centres on primary school (28.2%) and bachillerato (27.5%), with 14.3% holding no formal education. The segment skews toward cities and mid-sized towns outside the capital — La Libertad (18.7%), Santa Ana (12.1%), Sonsonate (10.9%); 71% urban.

What this suggests for marketing teams. The narratives describe women managing informal cash economies, running household purchasing, and making daily decisions with constrained but consistent budgets. Hypotheses worth testing: financial-access products and mobile money, family-security messaging, practical consumer goods, small-business tools framed around daily operational needs. The right message here almost certainly differs from what performs in formally employed urban audiences — and the right channel is unlikely to be digital-first.

Segment B — Physical labour and rural economy (30.0% · 6,004 personas)

The largest segment by persona count is almost exclusively male — 94.85% male — and concentrated in young ages, with the modal range 20–29. Occupations are dominated by grain and legume cultivation (1,058), building construction (636), vehicle maintenance and repair (207), land transport (185), and public security (164). Bachillerato is the modal education level (30.2%), followed by primary school (26.0%). This is the only segment with meaningful rural representation: 34.9% rural, against a sample average near 23%.

What this suggests for research and marketing teams. The narratives depict young men in physical, often outdoor work with limited formal infrastructure. Digital reach is likely lower than in the urban-services segment; connectivity gaps and distance from formal financial services recur as context. Hypotheses worth testing: offline and radio channels, messaging around practical tools and income stability, agricultural inputs, insurance framed around work-related risk, and community-oriented creative over individualistic aspiration.

Segment C — Urban services professionals (26.8% · 5,365 personas)

The most geographically concentrated and most gender-balanced segment. 71.3% come from San Salvador department — 32.3% from San Salvador Centro alone. It is the only segment with roughly equal male and female personas (50.2% male, 49.8% female). University education reaches 17.2% — five points above the dataset average — and 89.2% are urban. Top occupations reflect a more diverse services economy: food services (360), retail (309), domestic employment (293), construction (250), other services (250).

What this suggests for research and marketing teams. This is the profile most brands picture when they imagine an El Salvador consumer. It is real — but it represents roughly one in four personas and concentrates almost entirely in the capital. Hypotheses worth testing: digital channels, financial services, professional-development products, lifestyle messaging. The higher education presence and urban concentration make this the segment most likely to respond to digital-first creative and platform-mediated commerce.

Segment D — Food economy entrepreneurs (14.4% · 2,874 personas)

The smallest and most occupationally distinct segment. Bakery product manufacturing accounts for 847 personas; food services and street food for 753. Non-specialised food retail (349), market food stalls (249), and specialised food retail (229) follow. 76.4% are female. Education concentrates at bachillerato (30.2%) and primary school (25.5%). Unlike Segment C, this group is geographically spread — San Salvador (15.9%), Santa Ana (11.8%), La Libertad (10.6%), Sonsonate (9.5%) each contribute meaningfully.

What this suggests for research marketing teams. The narratives depict women running their own food businesses — managing suppliers, pricing products, serving regular customers, navigating informal supply chains. What they lack, in narrative context, is access to formal credit, business-management infrastructure, and supply-chain stability. Hypotheses worth testing: microfinance and working capital, ingredient and packaging supply, point-of-sale tools, and business-management messaging framed around operational efficiency.

Visual 5

Segment size and gender composition across the working sample

Persona count per segment
Segment A5,757
Segment B6,004
Segment C5,365
Segment D2,874
Gender composition
Segment A
98% F
2% M
Segment B
5% F
95% M
Segment C
50% F
50% M
Segment D
76% F
24% M

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Visual

Four segment profile cards

Segment A — Informal service workers

28.8% of sample · 5,757 personas
Gender98% female
GeographyLa Libertad, Santa Ana, Sonsonate
WorkDomestic work · street food · retail · garments
EducationPrimary / Bachillerato
ChannelLikely offline-first audience

Segment B — Physical labour economy

30.0% of sample · 6,004 personas
Gender95% male
GeographyRural & peri-urban
WorkAgriculture · construction · transport · security
EducationBachillerato / Primary
ChannelRadio and community channels

Segment C — Urban services

26.8% of sample · 5,365 personas
Gender50/50 male / female
GeographySan Salvador (71% of segment)
WorkFormal services · administration · commerce
EducationBachillerato + University (17%)
ChannelDigital-first, platform-ready

Segment D — Food & bakery economy

14.4% of sample · 2,874 personas
Gender76% female
GeographyMulti-department spread
WorkBakery · street food · market stalls · food retail
EducationBachillerato / Primary
ChannelMixed · community-rooted
Liorant analysis · 20,000-persona working sample · segments are exploratory hypotheses, not verified customer profiles

Visual

Audience segment profiles across five dimensions

Dimension (% of segment)Segment A
Informal service
Segment B
Physical labour
Segment C
Urban services
Segment D
Food economy
Female98%5%50%76%
Urban71%65%89%70%
University-educated9%7%17%11%
From San Salvador dept.13%12%71%16%
Rural29%35%11%30%

Each cell is shaded within its own row — the darkest cell marks the segment that scores highest on that dimension. Approximate values from the 20,000-persona working sample.

Source: Created by Liorant based on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Research methods: before synthetic personas and after

Most teams working on El Salvador — or any Central American market — begin with one of three starting points: senior stakeholder intuition, recycled global persona templates ("María, 35, middle-income mother"), or a formal research cycle that takes six to ten weeks and costs more than $10,000 before the first structured hypothesis appears. Here is what changes when a synthetic persona dataset enters the workflow.

Visual

Research timeline: traditional vs. AI-assisted

Traditional research flow
Week 1–2 Define research brief + design survey instrument
Week 3–4 Source panel, launch fieldwork, collect responses
Week 5–6 Clean data, run statistical analysis
Week 7–8 Write report → first structured hypothesis arrives
$10,000+ · 4–10 weeks · one hypothesis cycle per fieldwork round
AI-assisted with synthetic personas
Hour 1 Remote EDA on the full 148,000-row dataset
Hour 2–3 Embed persona texts, run segmentation
Hour 4 Profile segments, draft audience hypotheses
Hour 5+ Enter fieldwork with structured, testable questions
≈ Only cost of computing + $0 marginal cost · hours, not weeks · hypotheses iterable in real time
Synthetic personas compress hypothesis formation — not the fieldwork that follows.

Source: Created by Liorant. Cost and timeline ranges are Liorant estimates for the Central American market.

DimensionTraditional approachWith synthetic personas
Time to first hypothesis4–10 weeksHours
Cost to produce segments$10,000+Open-source tooling + compute only
Geographic resolutionCity-level or broad regionMunicipality-level, all 14 departments
Occupational granularityBroad categories119 CIIU Rev.4 classes
Narrative depthClosed survey responsesLong-form narratives, 11 life dimensions
Iteration speedNew fieldwork per hypothesisRe-run in hours
AI workflow readinessLimited — structured data onlyDirect: embeddings, semantic search, LLM-ready
Validation requirementSelf-validating (real respondents)Requires real-data validation before any decision

Source: Created by Liorant. Comparison reflects Liorant's applied use of the NVIDIA Nemotron-Personas-El-Salvador dataset.

The final row is the most important. Synthetic personas accelerate the hypothesis-formation phase of research, not the conclusion phase. Every segment, pattern, and campaign idea produced from this dataset requires validation against real customer data, CRM records, field interviews, or campaign results before any strategic decision rests on it. The competitive advantage is not in replacing fieldwork — it is in arriving at fieldwork with better-formed questions.

An AI research assistant built on the same dataset

The third phase built an interactive layer on top of the segmentation work. The precomputed narrative vectors and the clustered sample were loaded into a lightweight pipeline connecting a semantic-search layer to a small instruction-following language model — running entirely on a standard cloud GPU at no incremental cost.

A user types a marketing question in natural language. The system identifies the five persona narratives most semantically similar to the question, then passes them to the language model, which generates a structured response — audience hypothesis, campaign brief, segment summary, or message framework — grounded in the retrieved personas rather than the model's general training.

Workflow

How the AI marketing-research assistant works

Stage 1
User types a marketing question in plain language
Stage 2
Question encoded as a semantic vector by the same model used for the personas
Stage 3
Top 5 most semantically similar persona narratives retrieved
Stage 4
Small language model reads the 5 personas + the question together
Stage 5
Structured output: hypotheses · motivations · barriers · channels · claims to validate
Example output

Question: "Which personas might respond to a digital financial-education campaign?"

Motivations
income stability, family financial security
Barriers
trust in informal providers, limited digital experience
Channels
WhatsApp groups, community radio, in-person referrals
Validate with
customer interviews, CRM data, pilot campaign results
Runs on a standard cloud GPU · outputs are AI-generated hypotheses, not validated market research · always validate with real customers before acting

Source: Created by Liorant. Pipeline built on the NVIDIA Nemotron-Personas-El-Salvador dataset.

Live interface

Creating a research query and consulting in natural language

liorant · AI research assistant

The assistant in use. A team member loads the source material and types a question in plain language — here in Spanish — and the system returns a structured, source-grounded answer in the same language. No query syntax, no data export, no technical step in between: the research question and the consultation happen in one conversation.

Source: Liorant working interface, built on the NVIDIA Nemotron-Personas-El-Salvador workflow.

This is not market research. It is a hypothesis accelerator — a tool that helps a marketing team enter a brief with structured, testable assumptions rather than unvalidated intuitions. A browser-accessible interface built at the end of the pipeline makes it usable by non-technical teams, requiring no data literacy to operate.

How to reproduce this workflow

The full pipeline runs in Google Colab's free tier with no local installation. The technical complexity sits in the notebooks; the business team's input is the question, the product, and the audience context.

  • Run the exploratory analysis. The first notebook connects to the dataset files remotely, runs all descriptive statistics as database queries, and exports a 20,000-row working sample.
  • Save the working sample. Download the sample file (~60–80 MB) before moving on.
  • Run the segmentation analysis. The second notebook generates narrative embeddings, evaluates multiple segment counts, applies the four-segment solution, profiles all segments, and exports the segmented sample + embedding file.
  • Save two output files. The embedding file and the segmented sample are both needed for the demo.
  • Run the marketing-research demo. The third notebook loads the two files, builds the semantic-search layer, connects a small language model, and launches a browser interface.

Workflow

Five steps from dataset to working marketing interface

1
Run remote EDA
First notebook queries the full 148,000-row dataset via Hugging Face — no local download
2
Save the working sample
Download the 20,000-row sample file (~60–80 MB)
3
Run segmentation
Second notebook embeds persona texts and applies four-segment clustering
4
Save two output files
Embeddings file + clustered sample needed for the demo step
5
Launch the demo
Third notebook builds a browser interface — type any question, receive persona-grounded hypotheses
Google Colab free tier · no local setup · no API subscription required

Source: Created by Liorant. Workflow built on the NVIDIA Nemotron-Personas-El-Salvador dataset.

What this workflow cannot replace

Synthetic personas accelerate hypothesis formation; they do not substitute for real-world validation. The Nemotron dataset is a starting point for structured hypothesis generation — one of the most valuable phases in a research process, and one that marketing teams in Latin America consistently underfund. It is not a shortcut past the research that follows.

Visual

What synthetic personas do not cover

Real customer interviews

Only real people can confirm whether a motivation or barrier actually applies to them.

CRM & behavioural data

Purchase frequency, churn signals, and lifetime value require transaction records.

Campaign performance

Only live A/B tests reveal whether a message actually performs.

Cultural & local expertise

No dataset replaces the knowledge of someone who lives within the culture being studied.

Ethical oversight

AI-assisted research affecting real communities requires human review and documented validation.

NVIDIA-documented dataset limits

Under-represents Indigenous and Afrodescendant communities · may carry inherited gender-role assumptions · Salvadoran Spanish approximated by generative AI · religion not modelled.

Synthetic personas accelerate hypothesis formation. Validation with real people, real data, and real campaigns is always required before any strategic decision.

Source: Created by Liorant. Dataset limits per the official NVIDIA dataset card.

How Liorant helps research and marketing teams build this

Liorant builds and runs AI systems for marketing, operations, and legal teams across Spain, Colombia, and El Salvador. The workflow in this article represents a category of capability we call AI Customer Intelligence — combining structured data, narrative embeddings, segmentation, semantic search, and lightweight language models into repeatable research tools that marketing teams can operate themselves.

A typical engagement starts with a product launch, a regional expansion, or an audience question existing research has not answered. Liorant designs and deploys an AI-assisted audience-research workflow calibrated to the client's data environment — CRM records, campaign data, product usage, or a synthetic dataset as a starting scaffold — and delivers a working system in four to six weeks.

The deliverable is not a report. It is a running tool the marketing team operates themselves, backed by Liorant's AI engineering team for ongoing maintenance, iteration, and expansion. We report results in business terms: hypotheses generated, research cycles compressed, campaign-brief quality improved.

See how Liorant delivers working AI systems in four to six weeks →

Frequently asked questions

Is the Nemotron-Personas-El-Salvador dataset free to use?

Yes. NVIDIA published it on Hugging Face under a Creative Commons Attribution 4.0 International (CC-BY 4.0) licence, which permits use for any purpose — including commercial — provided appropriate credit is given to NVIDIA. Always verify the current terms on the dataset card before use.

Do the four segments represent real Salvadoran consumers?

No. The segments are patterns that emerged from an AI-driven analysis of a 20,000-persona sample drawn from a synthetic dataset. They represent plausible groupings worth investigating, not verified customer profiles. Every hypothesis they generate requires validation with real people before informing a decision.

Do we need a technical team to work with this data?

Not if you work with Liorant. All technical design, implementation, model selection, and operations sit with Liorant's engineering team. A marketing executive contributes the business question, the strategic context, and the validation priorities.

What is the difference between a synthetic persona and a buyer persona?

A buyer persona is a composite built from real customer research. A synthetic persona is a computationally generated profile representing plausible demographic and narrative diversity, not a specific real individual. Buyer personas summarise validated knowledge; synthetic personas explore the hypothesis space before that validation begins. The two are complementary, not interchangeable.

What is the first step for a company interested in this capability?

A 30-minute discovery session with Liorant. We review your existing data assets, identify the audience question most worth answering, and outline how an AI-assisted research workflow would be calibrated to your market, product, and team — without any upfront commitment.

Start building AI-assisted audience research

Teams that wait two to three months for fieldwork before forming a single testable hypothesis will fall behind teams that enter the field with sharper questions. Synthetic persona analysis does not shorten fieldwork — it improves what teams bring to it.

The Nemotron dataset, freely available, analysed in hours, and surfacing four distinct audience patterns in El Salvador's synthetic population, is an early demonstration of how that shift works in practice. The teams that act on it first — validating the most promising hypotheses with real customers — will build a structural advantage in market understanding.

Start with a free 30-minute AI discovery session.

We identify your highest-value audience-research opportunity and explain exactly how Liorant can help — no slides, no pitch.

Book your session →