Skip to content

[BUG]: StridedMemoryView.from_dlpack ignores DLTensor.byte_offset #2592

Description

@fedonman

Type of Bug

Silent Failure

Component

cuda.core

Describe the bug

The DLPack specification places a tensor's first element at data + byte_offset. StridedMemoryView.from_dlpack sets ptr from data alone, so a producer that reports an allocation base in data and expresses a slice in byte_offset yields a view whose ptr is short by exactly byte_offset bytes. Nothing raises, and every consumer built on ptr reads the wrong memory, including the __dlpack__ re-export, which writes ptr back out as data with byte_offset = 0 and so loses the offset permanently. from_any_interface is affected too, since it delegates to from_dlpack for a DLPack producer. The capsule-consuming helper in the same module, _smv_from_dlpack_capsule, does add byte_offset.

How to Reproduce

This producer describes src[8:] the way the spec allows, with the allocation base in data and the slice in byte_offset. It exits 0 when ptr is correct and 1 when it is not.

import ctypes

import numpy as np
from cuda.core.utils import StridedMemoryView

i64p = ctypes.POINTER(ctypes.c_int64)


class DLTensor(ctypes.Structure):
    _fields_ = [("data", ctypes.c_void_p), ("device_type", ctypes.c_int32), ("device_id", ctypes.c_int32),
                ("ndim", ctypes.c_int32), ("code", ctypes.c_uint8), ("bits", ctypes.c_uint8),
                ("lanes", ctypes.c_uint16), ("shape", i64p), ("strides", i64p), ("byte_offset", ctypes.c_uint64)]


class DLManagedTensorVersioned(ctypes.Structure):
    _fields_ = [("major", ctypes.c_uint32), ("minor", ctypes.c_uint32), ("manager_ctx", ctypes.c_void_p),
                ("deleter", ctypes.c_void_p), ("flags", ctypes.c_uint64), ("dl_tensor", DLTensor)]


ctypes.pythonapi.PyCapsule_New.argtypes = (ctypes.c_void_p, ctypes.c_char_p, ctypes.c_void_p)
ctypes.pythonapi.PyCapsule_New.restype = ctypes.py_object

src = np.arange(16, dtype=np.int32)
offset = 8 * src.itemsize
shape, strides = (ctypes.c_int64 * 1)(8), (ctypes.c_int64 * 1)(1)
t = DLManagedTensorVersioned(major=1, minor=0)
# Describe src[8:] the way the DLPack spec allows: allocation base in `data`,
# the slice expressed as `byte_offset`. First element is at data + byte_offset.
t.dl_tensor = DLTensor(data=src.ctypes.data, device_type=1, device_id=0, ndim=1, code=0, bits=32,
                       lanes=1, shape=shape, strides=strides, byte_offset=offset)


class Producer:
    def __dlpack_device__(self):
        return (1, 0)  # kDLCPU

    def __dlpack__(self, stream=None, max_version=None):
        return ctypes.pythonapi.PyCapsule_New(ctypes.addressof(t), b"dltensor_versioned", None)


view = StridedMemoryView.from_dlpack(Producer(), stream_ptr=-1)
print("expected ptr :", src.ctypes.data + offset)
print("view.ptr     :", view.ptr)
print("expected data:", src[8:])
print("view data    :", np.from_dlpack(view))
raise SystemExit(0 if view.ptr == src.ctypes.data + offset else 1)

Output:

expected ptr : 387455296
view.ptr     : 387455264
expected data: [ 8  9 10 11 12 13 14 15]
view data    : [0 1 2 3 4 5 6 7]

The process exits 1. Reproduced against a build of main at commit 757731a.

Expected behavior

ptr should be data + byte_offset, so the view begins at the tensor's first element, and the reproducer prints the producer's data and exits 0.

Metadata

Metadata

Assignees

Labels

triageNeeds the team's attention

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions