Each Spark platform provides storage utilities for file operations and credentials. These utilities are automatically injected into the runtime environment by the platform and are NOT imported as regular Python packages.
Utility Name: mssparkutils
How It's Available:
- Automatically injected into
__main__module at runtime - Access via:
import __main__; mssparkutils = __main__.mssparkutils - Also available via:
import notebookutils
Key Operations:
# File System Operations
mssparkutils.fs.ls(path) # List files
mssparkutils.fs.exists(path) # Check existence
mssparkutils.fs.cp(src, dst, overwrite) # Copy files
mssparkutils.fs.mv(src, dst) # Move files
mssparkutils.fs.rm(path, recurse) # Delete files
mssparkutils.fs.head(path, maxBytes) # Read file content
mssparkutils.fs.put(path, content, overwrite) # Write file content
# Credentials
mssparkutils.credentials.getToken(audience) # Get OAuth token
# Runtime Context
notebookutils.runtime.context.get("currentWorkspaceId") # Get workspace IDStorage Paths:
- Format:
abfss://<workspace-id>@onelake.dfs.fabric.microsoft.com/<lakehouse-id>/Files/... - OneLake uses ADLS Gen2 protocol (abfss)
When Used:
- File operations in notebooks
- Job deployment (uploading files to OneLake)
- Bootstrap shim (downloading framework wheels)
- Authentication token retrieval
Utility Name: mssparkutils
How It's Available:
- Automatically injected into runtime
- Import via:
from notebookutils import mssparkutils - Access via:
import __main__; mssparkutils = __main__.mssparkutils
Key Operations:
# File System Operations (same as Fabric)
mssparkutils.fs.ls(path)
mssparkutils.fs.exists(path)
mssparkutils.fs.cp(src, dst, overwrite)
mssparkutils.fs.mv(src, dst)
mssparkutils.fs.rm(path, recurse)
mssparkutils.fs.head(path, maxBytes)
mssparkutils.fs.put(path, content, overwrite)
# Credentials
mssparkutils.credentials.getToken(audience)
# Notebook Operations
mssparkutils.notebook.run(path, timeout, arguments)
mssparkutils.notebook.exit(value)Storage Paths:
- Workspace primary storage:
abfss://<container>@<storage-account>.dfs.core.windows.net/... - Uses Azure Data Lake Storage Gen2
When Used:
- File operations in Synapse notebooks
- Job deployment (uploading to workspace storage)
- Bootstrap shim (downloading wheels)
- Notebook chaining
Utility Name: dbutils
How It's Available:
- Automatically injected into
__main__module at runtime - Access via:
import __main__; dbutils = __main__.dbutils - Can also import:
from pyspark.dbutils import DBUtils(for initialization)
Key Operations:
# File System Operations (DBFS)
dbutils.fs.ls(path) # List files
dbutils.fs.head(path, maxBytes) # Read file content
dbutils.fs.put(path, content, overwrite) # Write file content
dbutils.fs.cp(src, dst, recurse) # Copy files
dbutils.fs.mv(src, dst) # Move files
dbutils.fs.rm(path, recurse) # Delete files
dbutils.fs.mkdirs(path) # Create directory
# Notebook Operations
dbutils.notebook.run(path, timeout, arguments)
dbutils.notebook.exit(value)
# Secrets
dbutils.secrets.get(scope, key) # Get secret value
dbutils.secrets.list(scope) # List secrets
# Authentication
dbutils.notebook.entry_point.getDbutils() # Get DB utils context
.notebook().getContext() # Get notebook context
.apiToken().get() # Get API tokenStorage Paths:
- DBFS:
dbfs:/...or/dbfs/... - External storage:
s3://...,wasbs://...,abfss://...
When Used:
- File operations in Databricks notebooks
- Job deployment (uploading to DBFS)
- Bootstrap shim (downloading wheels)
- Secrets management
- Notebook workflows
The framework automatically detects which utility is available:
def get_storage_utils():
"""Get platform-specific storage utilities"""
try:
import __main__
# Try mssparkutils (Fabric/Synapse)
mssparkutils = getattr(__main__, "mssparkutils", None)
if mssparkutils:
return mssparkutils
except:
pass
try:
# Try dbutils (Databricks)
import __main__
dbutils = getattr(__main__, "dbutils", None)
if dbutils:
return dbutils
except:
pass
raise RuntimeError("Could not find storage utils (mssparkutils or dbutils)")Platform services provide unified file operation methods:
class PlatformService:
def exists(self, path: str) -> bool:
# Uses mssparkutils or dbutils internally
pass
def read(self, path: str) -> str:
# Uses mssparkutils.fs.head or dbutils.fs.head
pass
def write(self, path: str, content: str) -> None:
# Uses mssparkutils.fs.put or dbutils.fs.put
pass
def list(self, path: str) -> List[str]:
# Uses mssparkutils.fs.ls or dbutils.fs.ls
passThe bootstrap shim uses storage utilities to install framework wheels:
# In bootstrap_shim.py (generated)
storage_utils = get_storage_utils()
# List wheel files
files = storage_utils.fs.ls(artifacts_path)
wheel_files = [f.path for f in files if f.name.endswith('.whl')]
# Install each wheel
for wheel_path in wheel_files:
subprocess.check_call([
sys.executable, "-m", "pip", "install",
wheel_path,
"--disable-pip-version-check"
])# At module level (Synapse)
try:
from notebookutils import mssparkutils
add_to_registry = True
except:
mssparkutils = None
add_to_registry = False# At method level (Fabric/Databricks)
def some_method(self):
import __main__
utils = getattr(__main__, "mssparkutils", None)
if not utils:
utils = getattr(__main__, "dbutils", None)
if utils:
# Use the utility
files = utils.fs.ls(path)# Cache utility reference
class PlatformService:
def __init__(self):
self._storage_utils = None
def _get_storage_utils(self):
if not self._storage_utils:
import __main__
self._storage_utils = getattr(__main__, "mssparkutils", None)
return self._storage_utilsWrong:
import mssparkutils # ❌ This will fail
import dbutils # ❌ This will failCorrect:
import __main__
mssparkutils = getattr(__main__, "mssparkutils", None) # ✅
dbutils = getattr(__main__, "dbutils", None) # ✅These utilities are only available when running on the platform:
- ✅ Available: In notebooks, Spark jobs, interactive clusters
- ❌ Not available: Local development, unit tests, IDE
For local testing, mock these utilities:
class MockStorageUtils:
class fs:
@staticmethod
def ls(path):
return [] # Mock implementation| Feature | Fabric | Synapse | Databricks |
|---|---|---|---|
| Utility Name | mssparkutils |
mssparkutils |
dbutils |
| Import Method | __main__ or notebookutils |
notebookutils |
__main__ |
| File System | OneLake (ADLS Gen2) | ADLS Gen2 | DBFS |
| Storage Prefix | abfss:// |
abfss:// |
dbfs:// |
| Secrets | Via credentials API | Via credentials API | dbutils.secrets |
| Token Auth | credentials.getToken() |
credentials.getToken() |
apiToken().get() |
Mock the utilities:
def test_something(mocker):
# Mock mssparkutils
mock_utils = mocker.MagicMock()
mock_utils.fs.ls.return_value = [...]
mocker.patch("__main__.mssparkutils", mock_utils)
# Test code
result = my_function()Run on actual platform or use platform service fixtures:
@pytest.fixture
def fabric_service(fabric_credentials):
"""Create real Fabric service for testing"""
from kindling.platform_fabric import FabricService
return FabricService(config, logger, credential)