ValidMind for development 3 — Integrate custom tests
Learn how to use ValidMind for your end-to-end documentation process with our series of four introductory notebooks. In this third notebook, supplement ValidMind tests with your own and include them as additional evidence in your documentation.
This notebook assumes that you already have a repository of custom made tests considered critical to include in your documentation. A custom test is any function that takes a set of inputs and parameters as arguments and returns one or more outputs:
The function can be as simple or as complex as you need it to be — it can use external libraries, make API calls, or do anything else that you can do in Python.
The only requirement is that the function signature and return values can be "understood" and handled by the ValidMind Library. As such, custom tests offer added flexibility by extending the default tests provided by ValidMind, enabling you to document any type of record (model) or use case.
For a more in-depth introduction to custom tests, refer to our Implement custom tests notebook.
Learn by doing
Our course tailor-made for developers new to ValidMind combines this series of notebooks with more a more in-depth introduction to the ValidMind Platform — Developer Fundamentals
Prerequisites
In order to integrate custom tests with your documentation with this notebook, you'll need to first have:
# Make sure the ValidMind Library is installed%pip install -q validmind# Load your model identifier credentials from an `.env` file%load_ext dotenv%dotenv .env# Or replace with your code snippetimport validmind as vmvm.init(# api_host="...",# api_key="...",# api_secret="...",# model="...", document="documentation",)
Note: you may need to restart the kernel to use updated packages.
2026-10-02 20:33:45,072 - INFO(validmind.api_client): 🎉 Connected to ValidMind!
📊 Model: [ValidMind Academy] Model development (ID: cmalgf3qi02ce199qm3rdkl46)
📁 Document Type: model_documentation
Import sample dataset
Next, we'll import the same public Bank Customer Churn Prediction dataset from Kaggle we used in the last notebook so that we have something to work with:
from validmind.datasets.classification import customer_churn as demo_datasetprint(f"Loaded demo dataset with: \n\n\t• Target column: '{demo_dataset.target_column}' \n\t• Class labels: {demo_dataset.class_labels}")raw_df = demo_dataset.load_data()
Loaded demo dataset with:
• Target column: 'Exited'
• Class labels: {'0': 'Did not exit', '1': 'Exited'}
We'll apply a simple rebalancing technique to the dataset before continuing:
import pandas as pdraw_copy_df = raw_df.sample(frac=1) # Create a copy of the raw dataset# Create a balanced dataset with the same number of exited and not exited customersexited_df = raw_copy_df.loc[raw_copy_df["Exited"] ==1]not_exited_df = raw_copy_df.loc[raw_copy_df["Exited"] ==0].sample(n=exited_df.shape[0])balanced_raw_df = pd.concat([exited_df, not_exited_df])balanced_raw_df = balanced_raw_df.sample(frac=1, random_state=42)
Remove highly correlated features
Let's also quickly remove highly correlated features from the dataset using the output from a ValidMind test.
As you learned previously, before we can run tests you'll need to initialize a ValidMind dataset object:
# Register new data and now 'balanced_raw_dataset' is the new dataset object of interestvm_balanced_raw_dataset = vm.init_dataset( dataset=balanced_raw_df, input_id="balanced_raw_dataset", target_column="Exited",)
With our balanced dataset initialized, we can then run our test and utilize the output to help us identify the features we want to remove:
# Run HighPearsonCorrelation test with our balanced dataset as input and return a result objectcorr_result = vm.tests.run_test( test_id="validmind.data_validation.HighPearsonCorrelation", params={"max_threshold": 0.3}, inputs={"dataset": vm_balanced_raw_dataset},)
❌ High Pearson Correlation
The High Pearson Correlation test evaluates pairwise linear relationships between variables to identify feature pairs with elevated correlation relative to a defined threshold. The result table lists the top reported correlations, showing each variable pair, its Pearson coefficient, and pass/fail status under a maximum threshold of 0.3. Reported coefficients range from -0.1978 to 0.3345, with one pair exceeding the threshold and the remaining listed pairs classified as passing.
Key insights:
One pair exceeds threshold: The pair (Age, Exited) has a Pearson correlation coefficient of 0.3345, which is above the configured threshold of 0.3 and is the only listed result marked as Fail.
All other reported pairs pass: The remaining nine reported correlations are below the threshold in absolute value, with coefficients between -0.1978 and 0.1631, and all are marked Pass.
Largest negative correlation is modest: The most negative reported relationship is (IsActiveMember, Exited) at -0.1978, which remains below the threshold in absolute magnitude.
Most listed relationships are weak: Several reported coefficients are close to zero, including (NumOfProducts, Exited) at -0.0614, (Tenure, IsActiveMember) at -0.0578, (Balance, HasCrCard) at -0.0402, (NumOfProducts, IsActiveMember) at 0.0374, (Tenure, EstimatedSalary) at 0.0360, and (HasCrCard, IsActiveMember) at -0.0276.
The reported correlation structure is limited to a single pair above the configured threshold, with (Age, Exited) representing the strongest observed linear relationship in the table. All other listed variable pairs exhibit correlations below 0.3 in absolute value, indicating that the reported top correlations are otherwise modest to weak. Overall, the result shows one flagged linear association alongside a broader set of passing pairwise relationships.
Parameters:
{
"max_threshold": 0.3
}
Tables
Columns
Coefficient
Pass/Fail
(Age, Exited)
0.3345
Fail
(IsActiveMember, Exited)
-0.1978
Pass
(Balance, NumOfProducts)
-0.1776
Pass
(Balance, Exited)
0.1631
Pass
(NumOfProducts, Exited)
-0.0614
Pass
(Tenure, IsActiveMember)
-0.0578
Pass
(Balance, HasCrCard)
-0.0402
Pass
(NumOfProducts, IsActiveMember)
0.0374
Pass
(Tenure, EstimatedSalary)
0.0360
Pass
(HasCrCard, IsActiveMember)
-0.0276
Pass
# From result object, extract table from `corr_result.tables`features_df = corr_result.tables[0].datafeatures_df
Columns
Coefficient
Pass/Fail
0
(Age, Exited)
0.3345
Fail
1
(IsActiveMember, Exited)
-0.1978
Pass
2
(Balance, NumOfProducts)
-0.1776
Pass
3
(Balance, Exited)
0.1631
Pass
4
(NumOfProducts, Exited)
-0.0614
Pass
5
(Tenure, IsActiveMember)
-0.0578
Pass
6
(Balance, HasCrCard)
-0.0402
Pass
7
(NumOfProducts, IsActiveMember)
0.0374
Pass
8
(Tenure, EstimatedSalary)
0.0360
Pass
9
(HasCrCard, IsActiveMember)
-0.0276
Pass
# Extract list of features that failed the testhigh_correlation_features = features_df[features_df["Pass/Fail"] =="Fail"]["Columns"].tolist()high_correlation_features
['(Age, Exited)']
# Extract feature names from the list of stringshigh_correlation_features = [feature.split(",")[0].strip("()") for feature in high_correlation_features]high_correlation_features
['Age']
We can then re-initialize the dataset with a different input_id and the highly correlated features removed and re-run the test for confirmation:
# Remove the highly correlated features from the datasetbalanced_raw_no_age_df = balanced_raw_df.drop(columns=high_correlation_features)# Re-initialize the dataset objectvm_raw_dataset_preprocessed = vm.init_dataset( dataset=balanced_raw_no_age_df, input_id="raw_dataset_preprocessed", target_column="Exited",)
# Re-run the test with the reduced feature setcorr_result = vm.tests.run_test( test_id="validmind.data_validation.HighPearsonCorrelation", params={"max_threshold": 0.3}, inputs={"dataset": vm_raw_dataset_preprocessed},)
✅ High Pearson Correlation
The High Pearson Correlation test evaluates pairwise linear relationships among features to identify potentially redundant variables or multicollinearity. The reported output lists the top 10 strongest Pearson correlations observed in the dataset, along with each pair’s coefficient and pass/fail status relative to the configured threshold of 0.3. All reported coefficients fall within a narrow range from -0.1978 to 0.1631, and every listed pair is marked as Pass. The largest-magnitude relationships in the output are between IsActiveMember and Exited, Balance and NumOfProducts, and Balance and Exited.
Key insights:
No correlations exceed threshold: All 10 reported feature pairs are below the configured absolute correlation threshold of 0.3. Every pair is therefore classified as Pass in the test output.
Strongest observed relationship is modest: The largest absolute coefficient is -0.1978 for the pair (IsActiveMember, Exited). This is materially below the threshold and indicates that the strongest reported linear association remains limited in magnitude.
Balance appears in several top pairs: Balance is included in four of the 10 reported correlations: with NumOfProducts (-0.1776), Exited (0.1631), HasCrCard (-0.0402), and EstimatedSalary (0.0225). Within the reported output, Balance is the most recurrent feature among the top-ranked pairwise relationships.
Top relationships include both positive and negative coefficients: The listed coefficients include negative values such as (IsActiveMember, Exited) at -0.1978 and (Balance, NumOfProducts) at -0.1776, as well as positive values such as (Balance, Exited) at 0.1631. This indicates that the strongest reported linear relationships are mixed in direction rather than concentrated in a single pattern.
The reported correlation structure shows no pairwise linear relationships above the configured threshold, with all top 10 correlations classified as passing. The highest observed absolute correlation remains below 0.2, indicating that the strongest reported associations are modest in magnitude. Within the listed results, Balance appears most frequently among the top pairs, but these relationships also remain below the test threshold.
Parameters:
{
"max_threshold": 0.3
}
Tables
Columns
Coefficient
Pass/Fail
(IsActiveMember, Exited)
-0.1978
Pass
(Balance, NumOfProducts)
-0.1776
Pass
(Balance, Exited)
0.1631
Pass
(NumOfProducts, Exited)
-0.0614
Pass
(Tenure, IsActiveMember)
-0.0578
Pass
(Balance, HasCrCard)
-0.0402
Pass
(NumOfProducts, IsActiveMember)
0.0374
Pass
(Tenure, EstimatedSalary)
0.0360
Pass
(HasCrCard, IsActiveMember)
-0.0276
Pass
(Balance, EstimatedSalary)
0.0225
Pass
Train the model
We'll then use ValidMind tests to train a simple logistic regression model on our prepared dataset:
# First encode the categorical features in our dataset with the highly correlated features removedbalanced_raw_no_age_df = pd.get_dummies( balanced_raw_no_age_df, columns=["Geography", "Gender"], drop_first=True)balanced_raw_no_age_df.head()
CreditScore
Tenure
Balance
NumOfProducts
HasCrCard
IsActiveMember
EstimatedSalary
Exited
Geography_Germany
Geography_Spain
Gender_Male
1848
484
8
0.00
2
1
0
186136.48
0
False
True
True
6621
850
9
92899.27
2
1
0
97465.89
0
False
False
False
570
776
2
169824.46
1
1
0
169291.70
0
False
False
False
6300
850
1
121874.89
1
0
0
6865.41
1
True
False
False
4670
598
9
0.00
1
1
0
68462.59
1
False
False
False
# Split the processed dataset into train and testfrom sklearn.model_selection import train_test_splittrain_df, test_df = train_test_split(balanced_raw_no_age_df, test_size=0.20)X_train = train_df.drop("Exited", axis=1)y_train = train_df["Exited"]X_test = test_df.drop("Exited", axis=1)y_test = test_df["Exited"]
Let's initialize the ValidMind Dataset and Model objects in preparation for assigning predictions to each dataset:
# Initialize the datasets into their own ValidMind dataset objectsvm_train_ds = vm.init_dataset( input_id="train_dataset_final", dataset=train_df, target_column="Exited",)vm_test_ds = vm.init_dataset( input_id="test_dataset_final", dataset=test_df, target_column="Exited",)# Initialize the ValidMind model objectvm_model = vm.init_model(log_reg, input_id="log_reg_model_v1")
Assign predictions
Once the model is registered, we'll assign predictions to the training and test datasets:
2026-10-02 20:33:59,687 - INFO(validmind.vm_models.dataset.utils): Running predict_proba()... This may take a while
2026-10-02 20:33:59,688 - INFO(validmind.vm_models.dataset.utils): Done running predict_proba()
2026-10-02 20:33:59,688 - INFO(validmind.vm_models.dataset.utils): Running predict()... This may take a while
2026-10-02 20:33:59,690 - INFO(validmind.vm_models.dataset.utils): Done running predict()
2026-10-02 20:33:59,692 - INFO(validmind.vm_models.dataset.utils): Running predict_proba()... This may take a while
2026-10-02 20:33:59,692 - INFO(validmind.vm_models.dataset.utils): Done running predict_proba()
2026-10-02 20:33:59,693 - INFO(validmind.vm_models.dataset.utils): Running predict()... This may take a while
2026-10-02 20:33:59,693 - INFO(validmind.vm_models.dataset.utils): Done running predict()
Implementing a custom inline test
With the set up out of the way, let's implement a custom inline test that calculates the confusion matrix for a binary classification model.
An inline test refers to a test written and executed within the same environment as the code being tested — in this case, right in this Jupyter Notebook — without requiring a separate test file or framework.
You'll note that the custom test function is just a regular Python function that can include and require any Python library as you see fit.
Create a confusion matrix plot
Let's first create a confusion matrix plot using the confusion_matrix function from the sklearn.metrics module:
import matplotlib.pyplot as pltfrom sklearn import metrics# Get the predicted classesy_pred = log_reg.predict(vm_test_ds.x)confusion_matrix = metrics.confusion_matrix(y_test, y_pred)cm_display = metrics.ConfusionMatrixDisplay( confusion_matrix=confusion_matrix, display_labels=[False, True])cm_display.plot()
Next, create a @vm.test wrapper that will allow you to create a reusable test. Note the following changes in the code below:
The function confusion_matrix takes two arguments dataset and model. This is a VMDataset and VMModel object respectively.
VMDataset objects allow you to access the dataset's true (target) values by accessing the .y attribute.
VMDataset objects allow you to access the predictions for a given record (model) by accessing the .y_pred() method.
The function docstring provides a description of what the test does. This will be displayed along with the result in this notebook as well as in the ValidMind Platform.
The function body calculates the confusion matrix using the sklearn.metrics.confusion_matrix function as we just did above.
The function then returns the ConfusionMatrixDisplay.figure_ object — this is important as the ValidMind Library expects the output of the custom test to be a plot or a table.
The @vm.test decorator is doing the work of creating a wrapper around the function that will allow it to be run by the ValidMind Library. It also registers the test so it can be found by the ID my_custom_tests.ConfusionMatrix.
@vm.test("my_custom_tests.ConfusionMatrix")def confusion_matrix(dataset, model):"""The confusion matrix is a table that is often used to describe the performance of a classification model on a set of data for which the true values are known. The confusion matrix is a 2x2 table that contains 4 values: - True Positive (TP): the number of correct positive predictions - True Negative (TN): the number of correct negative predictions - False Positive (FP): the number of incorrect positive predictions - False Negative (FN): the number of incorrect negative predictions The confusion matrix can be used to assess the holistic performance of a classification model by showing the accuracy, precision, recall, and F1 score of the model on a single figure. """ y_true = dataset.y y_pred = dataset.y_pred(model=model) confusion_matrix = metrics.confusion_matrix(y_true, y_pred) cm_display = metrics.ConfusionMatrixDisplay( confusion_matrix=confusion_matrix, display_labels=[False, True] ) cm_display.plot() plt.close() # close the plot to avoid displaying itreturn cm_display.figure_ # return the figure object itself
You can now run the newly created custom test on both the training and test datasets using the run_test() function:
# Training datasetresult = vm.tests.run_test("my_custom_tests.ConfusionMatrix:training_dataset", inputs={"model": vm_model, "dataset": vm_train_ds},)
Confusion Matrix Training Dataset
The Confusion Matrix test evaluates classification outcomes by comparing predicted labels against true labels on the training dataset. The result is presented as a 2x2 matrix with counts for correct negative predictions, incorrect positive predictions, incorrect negative predictions, and correct positive predictions. In this training sample, the matrix shows 829 true negatives, 474 false positives, 465 false negatives, and 817 true positives. The diagonal cells contain the correct classifications, while the off-diagonal cells represent misclassifications.
Key insights:
Correct classifications are balanced: The model records 829 true negatives and 817 true positives, indicating similar volumes of correct predictions across both classes.
Misclassification counts are also similar: False positives total 474 and false negatives total 465, showing that the two error types occur at nearly the same frequency in the training dataset.
Diagonal counts exceed off-diagonal counts: Both correct classification counts are higher than their corresponding error counts, with 829 exceeding 474 for the negative class and 817 exceeding 465 for the positive class.
Observed class outcomes are nearly even: Total actual negatives equal 1,303 and total actual positives equal 1,282, indicating a training sample with closely balanced observed class labels.
The training confusion matrix shows that correct predictions exceed incorrect predictions for both classes, with comparable performance across negative and positive outcomes. Error counts are closely aligned between false positives and false negatives, indicating no strong asymmetry in the observed misclassification pattern. The underlying training sample is also nearly balanced by actual class, which supports direct comparison of outcomes across the two classes.
Figures
# Test datasetresult = vm.tests.run_test("my_custom_tests.ConfusionMatrix:test_dataset", inputs={"model": vm_model, "dataset": vm_test_ds},)
Confusion Matrix Test Dataset
The ConfusionMatrix test evaluates classification performance by comparing predicted labels against true labels on the test dataset. The confusion matrix shows counts for true negatives, false positives, false negatives, and true positives across the two class labels. In this result, the four observed cell counts are 221 for true negatives, 92 for false positives, 120 for false negatives, and 214 for true positives.
Key insights:
Correct predictions exceed errors: The model records 221 true negatives and 214 true positives, compared with 92 false positives and 120 false negatives. Both classes therefore contain more correct than incorrect classifications.
Negative class is identified slightly better: Correct negative classifications total 221, while incorrect positive assignments for negative cases total 92. This corresponds to fewer errors in the negative class than in the positive class, where 120 true cases are classified as negative.
False negatives exceed false positives: The model produces 120 false negatives versus 92 false positives. Misclassification is therefore somewhat more concentrated in missed positive cases than in incorrect positive assignments.
The confusion matrix indicates that model predictions are predominantly concentrated on the diagonal, with correct classifications in both the negative and positive classes exceeding the corresponding misclassifications. Error counts are present in both directions, with false negatives occurring more often than false positives. Overall, the result reflects a reasonably balanced distribution of correct predictions across classes, alongside a modestly higher rate of missed positive cases.
Figures
Add parameters to custom tests
Custom tests can take parameters just like any other function. To demonstrate, let's modify the confusion_matrix function to take an additional parameter normalize that will allow you to normalize the confusion matrix:
@vm.test("my_custom_tests.ConfusionMatrix")def confusion_matrix(dataset, model, normalize=False):"""The confusion matrix is a table that is often used to describe the performance of a classification model on a set of data for which the true values are known. The confusion matrix is a 2x2 table that contains 4 values: - True Positive (TP): the number of correct positive predictions - True Negative (TN): the number of correct negative predictions - False Positive (FP): the number of incorrect positive predictions - False Negative (FN): the number of incorrect negative predictions The confusion matrix can be used to assess the holistic performance of a classification model by showing the accuracy, precision, recall, and F1 score of the model on a single figure. """ y_true = dataset.y y_pred = dataset.y_pred(model=model)if normalize: confusion_matrix = metrics.confusion_matrix(y_true, y_pred, normalize="all")else: confusion_matrix = metrics.confusion_matrix(y_true, y_pred) cm_display = metrics.ConfusionMatrixDisplay( confusion_matrix=confusion_matrix, display_labels=[False, True] ) cm_display.plot() plt.close() # close the plot to avoid displaying itreturn cm_display.figure_ # return the figure object itself
Pass parameters to custom tests
You can pass parameters to custom tests by providing a dictionary of parameters to the run_test() function.
The parameters will override any default parameters set in the custom test definition. Note that dataset and model are still passed as inputs.
Since these are VMDataset or VMModel inputs, they have a special meaning.
When declaring a dataset, model, datasets or models argument in a custom test function, the ValidMind Library will expect these get passed as inputs to run_test() or run_documentation_tests().
Re-running the confusion matrix with normalize=True and our testing dataset looks like this:
# Test dataset with normalize=Trueresult = vm.tests.run_test("my_custom_tests.ConfusionMatrix:test_dataset_normalized", inputs={"model": vm_model, "dataset": vm_test_ds}, params={"normalize": True})
Confusion Matrix Test Dataset Normalized
The ConfusionMatrix test evaluates classification performance by comparing predicted labels with true labels, and this result presents the normalized confusion matrix for the test dataset. The matrix shows the proportion of observations in each outcome category rather than raw counts. The displayed cell values are 0.34 for true negatives, 0.14 for false positives, 0.19 for false negatives, and 0.33 for true positives.
Key insights:
Correct classifications dominate: The diagonal cells account for 0.34 true negatives and 0.33 true positives, for a combined normalized share of 0.67.
Error mass is lower than correct mass: The off-diagonal cells sum to 0.33, comprising 0.14 false positives and 0.19 false negatives.
False negatives exceed false positives: The false negative cell is 0.19 versus 0.14 for false positives, indicating more missed positive cases than incorrect positive assignments.
Balanced correct-class recognition: The two correct-class cells are similar in magnitude, with 0.34 for negatives and 0.33 for positives.
The normalized confusion matrix indicates that most observations fall into correct classification cells, with roughly equal contributions from true negatives and true positives. Misclassifications are present at a lower combined share, and these errors are more concentrated in false negatives than false positives. Overall, the result reflects stronger mass on the diagonal than on the off-diagonal cells, with a relatively balanced distribution across the two correctly predicted classes.
Parameters:
{
"normalize": true
}
Figures
Log the confusion matrix results
As we learned in 2 — Start the model development process under Documenting results > Run and log an individual tests, you can log any result to the ValidMind Platform with the .log() method of the result object, allowing you to then add the result to the documentation.
You can now do the same for the confusion matrix results:
result.log()
2026-10-02 20:34:20,416 - INFO(validmind.vm_models.result.result): Test driven block with result_id my_custom_tests.ConfusionMatrix:test_dataset_normalized does not exist in model's document
Note the output returned indicating that a test-driven block doesn't currently exist in your documentation for this particular test ID.
That's expected, as when we run individual tests the results logged need to be manually added to your documentation within the ValidMind Platform.
Using external test providers
Creating inline custom tests with a function is a great way to customize your documentation. However, sometimes you may want to reuse the same set of tests across multiple records (models) and share them with others in your organization. In this case, you can create an external custom test provider that will allow you to load custom tests from a local folder or a Git repository.
In this section you will learn how to declare a local filesystem test provider that allows loading tests from a local folder following these high level steps:
Create a folder of custom tests from existing inline tests (tests that exist in your active Jupyter Notebook)
Let's start by creating a new folder that will contain reusable custom tests from your existing inline tests.
The following code snippet will create a new my_tests directory in the current working directory if it doesn't exist:
tests_folder ="my_tests"import os# create tests folderos.makedirs(tests_folder, exist_ok=True)# remove existing testsfor f in os.listdir(tests_folder):# remove files and pycacheif f.endswith(".py") or f =="__pycache__": os.system(f"rm -rf {tests_folder}/{f}")
After running the command above, confirm that a new my_tests directory was created successfully. For example:
~/notebooks/tutorials/development/my_tests/
Save an inline test
The @vm.test decorator we used in Implementing a custom inline test above to register one-off custom tests also includes a convenience method on the function object that allows you to simply call <func_name>.save() to save the test to a Python file at a specified path.
While save() will get you started by creating the file and saving the function code with the correct name, it won't automatically include any imports, or other functions or variables, outside of the functions that are needed for the test to run. To solve this, pass in an optional imports argument ensuring necessary imports are added to the file.
The confusion_matrix test requires the following additional imports:
import matplotlib.pyplot as pltfrom sklearn import metrics
Let's pass these imports to the save() method to ensure they are included in the file with the following command:
confusion_matrix.save(# Save it to the custom tests folder we created tests_folder, imports=["import matplotlib.pyplot as plt", "from sklearn import metrics"],)
2026-10-02 20:34:20,891 - INFO(validmind.tests.decorator): Saved to /home/runner/work/documentation/documentation/site/notebooks/EXECUTED/development/my_tests/ConfusionMatrix.py!Be sure to add any necessary imports to the top of the file.
2026-10-02 20:34:20,892 - INFO(validmind.tests.decorator): This metric can be run with the ID: <test_provider_namespace>.ConfusionMatrix
# Saved from __main__.confusion_matrix
# Original Test ID: my_custom_tests.ConfusionMatrix
# New Test ID: <test_provider_namespace>.ConfusionMatrix
Now that your my_tests folder has a sample custom test, let's initialize a test provider that will tell the ValidMind Library where to find your custom tests:
ValidMind offers out-of-the-box test providers for local tests (tests in a folder) or a Github provider for tests in a Github repository.
You can also create your own test provider by creating a class that has a load_test method that takes a test ID and returns the test function matching that ID.
For most use cases, using a LocalTestProvider that allows you to load custom tests from a designated directory should be sufficient.
The most important attribute for a test provider is its namespace. This is a string that will be used to prefix test IDs in model documentation. This allows you to have multiple test providers with tests that can even share the same ID, but are distinguished by their namespace.
Let's go ahead and load the custom tests from our my_tests directory:
from validmind.tests import LocalTestProvider# initialize the test provider with the tests folder we created earliermy_test_provider = LocalTestProvider(tests_folder)vm.tests.register_test_provider( namespace="my_test_provider", test_provider=my_test_provider,)# `my_test_provider.load_test()` will be called for any test ID that starts with `my_test_provider`# e.g. `my_test_provider.ConfusionMatrix` will look for a function named `ConfusionMatrix` in `my_tests/ConfusionMatrix.py` file
Run test provider tests
Now that we've set up the test provider, we can run any test that's located in the tests folder by using the run_test() method as with any other test:
For tests that reside in a test provider directory, the test ID will be the namespace specified when registering the provider, followed by the path to the test file relative to the tests folder.
For example, the Confusion Matrix test we created earlier will have the test ID my_test_provider.ConfusionMatrix. You could organize the tests in subfolders, say classification and regression, and the test ID for the Confusion Matrix test would then be my_test_provider.classification.ConfusionMatrix.
Let's go ahead and re-run the confusion matrix test with our testing dataset by using the test ID my_test_provider.ConfusionMatrix. This should load the test from the test provider and run it as before.
result = vm.tests.run_test("my_test_provider.ConfusionMatrix", inputs={"model": vm_model, "dataset": vm_test_ds}, params={"normalize": True},)result.log()
Confusion Matrix
The Confusion Matrix test evaluates classification performance by comparing predicted labels with true labels across the four outcome types: true negatives, false positives, false negatives, and true positives. The displayed matrix is normalized and presents the share of observations in each cell of the 2x2 classification table. The largest cells are the true negative cell at 0.34 and the true positive cell at 0.33, while the off-diagonal error cells are 0.14 for false positives and 0.19 for false negatives.
Key insights:
Correct classifications dominate: The diagonal cells sum to 0.67, with 0.34 true negatives and 0.33 true positives, indicating that most observations were classified correctly in the normalized matrix.
False negatives exceed false positives: The false negative cell is 0.19 compared with 0.14 for false positives, showing a higher share of missed positive cases than incorrectly flagged negative cases.
Balanced correct-class capture: The two correct classification cells are closely aligned at 0.34 and 0.33, indicating similar normalized shares of correctly identified negative and positive cases.
Error mass remains material: The off-diagonal cells sum to 0.33, showing that one-third of normalized outcomes fall into misclassification categories.
The normalized confusion matrix shows that correct classifications account for the majority of outcomes, with nearly equal shares of true negatives and true positives. Misclassifications remain concentrated in both error types, with false negatives occurring more frequently than false positives. Overall, the result reflects stronger weight on the diagonal than the off-diagonal, alongside an observable asymmetry in the model’s error distribution.
Parameters:
{
"normalize": true
}
Figures
2026-10-02 20:34:27,726 - INFO(validmind.vm_models.result.result): Test driven block with result_id my_test_provider.ConfusionMatrix does not exist in model's document
Again, note the output returned indicating that a test-driven block doesn't currently exist in your model's documentation for this particular test ID.
That's expected, as when we run individual tests the results logged need to be manually added to your documentation within the ValidMind Platform.
Add test results to documentation
With our custom tests run and results logged to the ValidMind Platform, let's head to the model we connected to at the beginning of this notebook and insert our test results into the documentation (Learn more:Work with test results):
From the Inventory in the ValidMind Platform, go to the model you connected to earlier.
In the left sidebar that appears for your model, click Development under Documents.
Locate the Data Preparation section and click on 3.2. Model Evaluation to expand that section.
Hover under the Pearson Correlation Matrix content block until a horizontal dashed line with a + button appears, indicating that you can insert a new block.
Click + and then select Test-Driven Block under FROM LIBRARY:
Click on Custom under TEST-DRIVEN in the left sidebar.
Select the two custom ConfusionMatrix tests you logged above:
Finally, click Insert 2 Test Results to Document to add the test results to the documentation.
Confirm that the two individual results for the confusion matrix tests have been correctly inserted into section 3.2. Model Evaluation of the documentation.
In summary
In this third notebook, you learned how to:
Next steps
Finalize testing and documentation
Now that you're proficient at using the ValidMind Library to run and log tests, let's put the last pieces in place to prepare our fully documented sample model for review: 4 — Finalize testing and documentation