Skip to content

Analysis Config File Reference

Chen Wang edited this page Oct 2, 2026 · 1 revision

Each SMILE analysis is defined by a JSON config file in www/routes/analyses/ of standalone-smm-smile. SMILE reads these files at startup to build the analysis page, list the analysis on the home page and in the Analytics Tools menu, and route submissions to the right analysis container. For the end-to-end workflow, see Adding a New Analysis to SMILE.

The analysis config files

Complete examples: sentiment.json (simple), autophrase.json (sliders), networkx.json, preprocessing.json, topic.json, NER.json, and classification.json (multi-step, with a custom script).

How the fields map to the page

Config fields on the analysis form

Config fields on the results section

Top-level fields

Field Required Description
path Yes URL of the analysis page. {"path": "sentiment"} serves the page at /sentiment. Don't include a slash.
title Yes Heading of the analysis page, and its name in menus.
imgURL Yes Icon shown for the analysis on the home page, relative to www/public/, for example bootstrap/img/logo/NLP/SA.png.
introduction Yes Short description on the home page, as an array of strings that SMILE joins with spaces. HTML is allowed.
wiki No Link for the home page's read more... button.
result_path Yes Where outputs are stored, relative to the user's storage folder. Follow the pattern /{category}/{path}/, for example /NLP/sentiment/. Existing categories are NLP (natural language processing), ML (machine learning), NW (network analysis) and GraphQL (collected social media data); you can add your own. At run time, user jdoe's sentiment results go to <bucket>/jdoe/NLP/sentiment/<run id>/.
args Yes List of the form input ids to send to the algorithm. Your algorithm receives each one under the same name in params. Use [] if there are no inputs.
results Yes The output files and how to display them. See Results.
get Yes, for a page The form. See Form. Configs without get (such as the classification split, train and predict steps) define a run endpoint but no page of their own.
post Yes, to run Where and how the analysis runs. See Run settings.
put No Like post, but the request first uploads a file (used by classification training). Contact the developers before using it.
custom_script No Path to an extra JavaScript file for complex interfaces, for example bootstrap/js/customized/analyses/view_classification.js. Contact the developers before using it.

How args connects everything

An argument named tagger must appear in three places with the same name:

{
  "args": ["tagger"],
  "get": {
    "containers": [
      {
        "container-name": "",
        "container-label-name": "Tagger",
        "input": { "type": "select", "name": "tagger", "id": "tagger", "options": [ ... ] }
      }
    ]
  }
}
def algorithm(df, params):
    tagger = params['tagger']   # whatever the user selected

Results

results is a list with one entry per output file. Each acronym must match a key in the dictionary returned by algorithm().

Field Required Description
acronym Yes Output name, without file extension. Must be unique within the analysis.
name Yes Human-readable name shown on the page and in the download list.
download Yes true to list the file under Download.
img Yes true if the output is an HTML visualization to embed in the page.
preview Yes true to show the output (a CSV) as a table preview.
dataTable No With preview: true, set true to add search, paging and column sorting to the table.
wordtree No true to render the output (one phrase per line) as a Google word tree.
config Yes true if this output is the run's configuration file. Every analysis should include one entry like {"acronym": "config", "name": "configuration", "download": true, "img": false, "preview": false, "config": true}; the analysis container saves the run parameters under that name automatically.

Form (get)

get.containers

A list of form rows, displayed in order below the dataset picker.

Field Required Description
container-name Yes HTML id of the row. Can be an empty string.
container-label-name Yes Label shown on the left of the row.
container-classname No CSS class of the row. Custom scripts use it to show or hide groups of rows.
input Yes The input control. Its type must be one of the types below. To add another type, contact the developers.

Every input needs a name and a unique id. The id is what goes in args.

select, a dropdown:

Field Required Description
type Yes "select"
name, id Yes Name and unique id
options Yes List of {"value": ..., "label": ...}. By convention, the first option is {"value": "Please Select...", "label": "Please Select..."}.

range, a slider:

Field Required Description
type Yes "range"
name, id Yes Name and unique id
min, max, value Yes Minimum, maximum and default value
step No Step size
output_id Yes id of the element that displays the current value, for example "rangeNode"
onchange Yes Updates that display, for example "rangeNode.value=value"
{
  "container-name": "minimum-support",
  "container-label-name": "Minimum Support",
  "input": {
    "type": "range", "min": "1", "max": "100", "value": "3",
    "name": "minSup", "id": "minSup",
    "onchange": "rangeNode.value=value", "output_id": "rangeNode"
  }
}

text, a text box: type: "text", name, id.

radio, a radio button: type: "radio", name, id, value, label, and optionally checked.

file-upload, an upload button:

Field Required Description
type Yes "file-upload"
name, id Yes Name and unique id
displayname Yes id of the element that shows the chosen file's name
style No CSS for the underlying file input, for example "width: 0.1px; height: 0.1px; opacity: 0; overflow: hidden; position: absolute; z-index: -1;" to hide the browser's default control

uid, a text box for an identification code that links the steps of a multi-step analysis such as classification: type: "uid", name, id.

get.buttonGroups

The buttons at the end of the form. For a standard analysis, use the default Submit/Clear pair:

"buttonGroups": [
  { "id": "clear",  "class": "btn btn-primary", "value": "Clear",  "style": "margin: auto 5px;", "onclick": "customized_reset();" },
  { "id": "submit", "class": "btn btn-danger",  "value": "Submit", "style": "margin: auto 5px;" }
]

The button with "id": "submit" triggers the standard submit flow. An analysis with several steps (for example split, train and predict) needs one button per step, each with its own onclick function defined in a custom_script. Contact the developers for details.

get.citation

References to show when the user picks a particular option, in addition to SMILE's own citations. All existing configs include this field, and the page's citation script expects it to be present.

Field Required Description
trigger_id Yes id of the input to watch, for example "algorithm"
content Yes List of {"text": [...], "condition": ...}. When the watched input's value equals condition, each string in text is shown as a reference. HTML links are allowed. Use "condition": "!" to show a reference for any selection.
"citation": {
  "trigger_id": "algorithm",
  "content": [
    {
      "text": ["Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. ICWSM-14."],
      "condition": "vader"
    }
  ]
}

Run settings (post)

Field Required Description
rabbitmq_queue Yes Name of the RabbitMQ queue the analysis container listens on. Must equal the container's QUEUE_NAME environment variable.
lambda_config For real-time mode Its presence enables real-time mode: SMILE waits for the result and displays it on the page. Set it to {"aws_lambda_function": "<name>"}. The name is only used when SMILE runs analyses outside the local containers (LOCAL_ALGORITHM=false), so locally any descriptive name works.
batch_config For background mode Its presence enables background mode: the analysis runs in the background and the user is emailed when it finishes. batch_action and batch_script form the command run inside the container, for example "python3" and "/scripts/batch_function.py".
cutoff No When both modes are enabled, datasets with at most this many rows run in real time and larger ones in the background. Defaults to 5000.

With only lambda_config, an analysis always runs in real time. With only batch_config, it always runs in the background.

"post": {
  "cutoff": 10000,
  "lambda_config": { "aws_lambda_function": "sentiment_analysis" },
  "batch_config": { "batch_action": "python3", "batch_script": "/scripts/batch_function.py" },
  "rabbitmq_queue": "sentiment_analysis"
}

Clone this wiki locally