- The project is an autoEDA framework with the rich user experience.
- Any non-standard plots appearing in the report is a courtesy of authors mind.
- It is also capable of processing some particular bioinformatics data formats (protein sequences, DNA, SMILES strings), because I myself am too lazy to do explicit exploratorials by hand every time.
- Some of the captions might be hard-coded.
- The report header is designed to add some amount of exaggeration.
- The plots are generated 'as is', after that they are processed with LLMs. LLMs might hallucinate. I also might hallucinate during plots coding. This means that any results might not correctly represent the input data. Use at your own risk.
First, you need to install it as a package
git clone -b dev https://github.com/latticetower/storybear.git
pip install -e storybear
After that, storytell command should be available via shell.
To load sample data file, run the command
storygen smth.csv
as a result, there will appear file named smth.csv in the current directory.
Select your csv file (in the example below it is named smth.csv and located in the current directory) with tabular data and run
storytell smth.csv --tempdir imgdir
In the example above, imgdir is a path to a directory where the plot files, including temporary ones, will be located.
This produces the report in docx format.
There is also an option to run demo with gradio. This can be done by running
gradio app.py
The names of the classes representing each particular part of the pipeline are selected based on their function.
- The pipeline accepts .csv file with columns of different types as an input
- The columns are processed by the class which I named
DataGal. As a result of processing, I have a lot of plots produced from dataset columns (let's call each of them P) in the temporary folder. - For each P and corresponding statistical info (represented as a dictionary of key and value pairs, differs for different plotter classes) I use class named
Captionist- to generate text caption C with LLM. - Each pair of (plot P, text caption C) is processed by
Foodieclass - which is also powered by LLM and returns ranking R. It could have been named Critique, but I've decided to keep it simple in case if I'll decide to add other filtering steps in the pipeline and call them Critique. - Next class
Secretaryin the pipeline accepts all the triplets (plot P, text caption C, ranking R) and pick top N (N is a fixed parameter) samples with the highest rankings R. Technically, it is little filter in the beginning ofEditorexecution - I didn't included it in the scheme image. - The results of the previous step are processed by
Editorclass.Editoruses all the data from the previous step, to summarize them to report lead L (which is an analogue to lead in a news article) and a catchy header H. Header is a more or less exaggerrated depending on some parameter. - The top N triplets (P, C, R) from step 4 with the report lead L, header H from the step 5 are collected and given to the class
Junior. Basically, in this step the LLM does routine work: look at the editor's idea of the paper and reorder all the selected triplets (P, C, R) to make the overall story look more convincing. - The data produced by step 7 are given to the
Artistclass, which modifies the plots in place using its own artistic vision with LLM and keeps everything else as it was. - The last class is
Typography- it collects the data from the previous step and converts the report (for simplicity, made withpython-docxpackage).
- add other methods of reporting, i.e., probably replace docx with pdf generation
- Remove columns with "_id", "Id" and "identifier" from the consideration
- add default text gen (both with and without LLMs, no data)
- remove very similar plots based on their descriptions
- don't save plots which were filtered by description
- don't build plots for the highly correlated columns
- For string columns: draw length distributions
- Proteins: compute embeddings with esm2 8m + draw scatterplots
- Proteins: add color
- SMILES: chemberta
- add some sort of log to be able to understand why some of the columns were not processed
- add example with smiles and/or RNA/DNA
- add new types of columns
- add Modal code to main repo
- Add custom flowerplot, heatmap, and something fun.
- https://github.com/py-pdf/fpdf2 package for pdf report creation
- https://arxiv.org/abs/2605.14163 possible candidate method for pipeline inprovement
- https://arxiv.org/abs/2508.16757 paper on reranking strategies. I use basic and slow approach at the moment - pair reranking of plot descriptions, scoring based on this reranking, selection of top N plots (N=5).
- https://huggingface.co/black-forest-labs/FLUX.2-klein-4B in local inference, for image styling
- https://huggingface.co/openbmb/MiniCPM-V-4.6 in main pipeline, for caption, header, lead generation
- https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 source of text embeddings for plots filtering and creation
- https://huggingface.co/facebook/esm2_t6_8M_UR50D source of protein embeddings for corresponding columns (if any). My favourite model!
- https://huggingface.co/RaphaelMourad/Mistral-DNA-v1-138M-bacteria source of DNA/RNA embeddings for corresponding columns (if any). I don't use DNA/RNA models, so I took the first small one
- https://huggingface.co/DeepChem/ChemBERTa-10M-MLM source of molecular embeddings for corresponding columns with SMILES strings
- @latticetower Tatiana Malygina
- @nofate Michael Gamov

