[Beowulf] InfiniBand VL15 error
Many of your questions may have already been answered in earlier discussions or in the FAQ. The search results page will indicate current discussions as well as past list serves, articles, and papers.
Prentice Bisbal prentice at ias.eduTue Dec 2 07:24:15 PST 2008
- Previous message: [Beowulf] Personal Introduction & First Beowulf Cluster Question
- Next message: [Beowulf] InfiniBand VL15 error
- Messages sorted by: [ date ] [ thread ] [ subject ] [ author ]
I'm getting this error when I run ibchecknet on my cluster: #warn: counter VL15Dropped = 476 (threshold 100) lid 1 port 1 Error check on lid 1 (aurora HCA-1) port 1: FAILED I've googled around this morning, but haven't found anything helpful. Most of the hits turn up code with the phrase "VL15Dropped", but nothing explaining what this error means, what causes it, or how to fix it. After clearing the counters with 'perfquery -r', the VL15Dropped count starts increasing from zero almost immediately. Any ideas what this error represents or how to fix? Could it be a bad cable? -- Prentice
- Previous message: [Beowulf] Personal Introduction & First Beowulf Cluster Question
- Next message: [Beowulf] InfiniBand VL15 error
- Messages sorted by: [ date ] [ thread ] [ subject ] [ author ]
More information about the Beowulf mailing list
